Adaptive Mass-Segmented KV Compression for Long-Context Reasoning
Abstract
Long-context autoregressive reasoning in large language models is severely bottlenecked by key-value (KV) cache memory, making decoding-time compression essential. Existing methods typically rely on global or per-head token-wise selection. Under tight memory budgets and repeated compression events, however, these policies tend to over-focus on local attention spikes, causing Region Wipe-out: the severe depletion of contiguous spans of intermediate reasoning from the cache. This fragmentation disrupts reasoning, leading to Problem Drifting and Repetition Collapse. To address this, we propose Adaptive Mass-Segmented (AMS), a scorer-agnostic KV compression framework based on an "allocate-then-score" paradigm. By decoupling budget allocation from token scoring, AMS acts as a plug-in that shifts attention from a strictly micro-level retention metric to a macro-level spatial partitioning tool. Specifically, AMS derives an attention-based quality-mass distribution to partition each head's cache into adaptive segments, enforcing region-wise quotas with minimum-keep guarantees. Furthermore, an exponential moving average (EMA) credit mechanism makes this allocation history-aware, smoothing cache evolution across repeated recompressions. Experiments on MATH500, AIME24, AIME25, and GSM8K using 7B and 32B backbones show AMS improves pass@1 accuracy by up to 20.0 points over corresponding base scorers. AMS demonstrates generalization across diverse long-context workloads, including code completion, open-domain QA, and sparse retrieval. AMS seamlessly integrates with token-level scorers, consistently improves their performance, sometimes surpasses uncompressed full-KV decoding, and incurs negligible practical decoding-time overhead.