Coverage-Verified Sparse Attention: Closed-Loop Quality Control for Long-Context LLM Decoding
Abstract
Sparse attention accelerates long-context decoding by reading a subset of the KV cache, but existing methods are open-loop: the kernel commits to its selection with no signal indicating whether enough attention mass was preserved and no recovery path when it was not. The needed signal already exists inside the kernel. The fraction of softmax mass on the selected tokens, which we call coverage, satisfies an exact algebraic equality with per-step output error and serves a dual role as the acceptance criterion and the interpolation weight for recovery. Coverage-Verified Sparse Attention (CVSA) turns coverage into a training-free, closed-loop verify-then-correct decode loop that accepts high-coverage drafts and recovers only the heads that fall short, with no sampling, no draft model, and no speculative-decoding infrastructure. Turning verification on matters far more than tuning its threshold, and matched-budget comparisons confirm the gain is structural. On LongBench and RULER at 7B and 70B, CVSA closes the quality gap to dense attention to within statistical noise, with measured throughput reaching 6.21x at 128K context.