Coverage, Not Ranking: Re-examining KV-Staleness Signals for Block-Diffusion Language Models
Abstract
Block-diffusion language models commit KV-cache entries at block finalisation, so the operationally relevant staleness horizon runs from a token's crossing time to the freezing of its block. We show that the standard drift label measured over this horizon is confounded: it grows mechanically with how long an entry dwells in cache before the block freezes. On LLaDA-8B-Instruct (200 prompts, 25,600 tokens) dwell time alone predicts block-finalisation drift at rho = +0.610, better than any zero-cost signal; projecting it out costs entropy 81% of its correlation (+0.515 -> +0.098) while distance-to-nearest-unresolved-token loses 4% (-0.213 -> -0.203). We then ask whether the surviving signal changes a caching decision, and find it does not -- but for an instructive reason. Comparing refresh policies under an exact cache simulation, every static-scoring policy loses to uniform random refresh, including an oracle with perfect knowledge of true final drift (-0.176, p < 0.0001), while every time-varying policy ties it -- even a dynamic oracle with exact per-step staleness. A replay of the refresh schedules explains why: static scoring starves the same entries indefinitely, and adding a staleness term to a rotation schedule more than doubles the worst starvation gap. The binding constraint on block-diffusion KV refresh is coverage, not ranking. We release the harness, 10,688 labelled token records with 32-layer drift traces, and all analysis code.