Certified Continuation and the Contraction Spectrum: Output-Certified Head–Block Routing for Long-Context Decoding
Abstract
Decoding from a long-context language model reads the entire key-value (KV) cache for every generated token, and on memory-bandwidth-limited devices this read can dominate latency. Sparse-attention methods reduce the read by selecting a subset of the cache, but heuristic selection can miss entries a query needs. Our work studies certified continuation, which maintains an upper bound on the omitted attention mass together with a current-query lower bound on the retained mass. A continuation check carries the certificate across decode steps using query-drift envelopes, and falls back to fresh certification or dense attention when the check is inconclusive; soundness is stated in exact arithmetic. On frozen Qwen2.5-0.5B attention operands at 32K context, the corrected all-head report records zero tested certificate violations and retained-row reductions of 2.8-7.0x per head (1.39-2.56x for the GQA union), computed as reciprocals of mean retained fractions. The certified masks also reveal a descriptive contraction spectrum across heads and a 2.0-2.7x union head-block work-count overhead, and the mass-derived output bound exceeds the measured local output error by 11-13x (ratios of medians). Together, these results establish a sound, fail-closed guarantee for sparse decoding on a device-class model and locate the next gains in per-head routing and output-aware certification.