The Entropy–Drift Correlation Is Not One Quantity
Abstract
A growing line of work on KV-cache reuse for diffusion language models rests on a reported association between a token's decoding entropy and the subsequent drift of its cached key/value tensors. The association is normally quoted as a single coefficient. We show that it is not a single quantity. Computing "the entropy-drift correlation" requires at least five choices -- which positions enter the average, which layers, token- or step-level granularity, which drift horizon, and whether EOS decodes are kept -- and published descriptions generally state none of them. Sweeping 28 defensible combinations on one fixed dataset of 200 prompts, the coefficient ranges from -0.215 to +0.624: a span of 0.839 that crosses zero. The sign of the effect is therefore a free parameter of unstated analysis choices. We emphasise what this does and does not show. It is not evidence that any published value is wrong: the largest previously reported figure we are aware of, 0.644 (Cheong et al., 2026), sits at the top of the range our own data spans and is reachable from our data by choosing early layers with a block-final horizon (+0.624). The published result is recoverable, not refuted. What fails is comparability -- two papers reporting this coefficient may not be measuring the same thing, and a practitioner cannot tell. We close with the minimal reporting standard that would make the quantity identified.