AsdaKV: Attention-Overlap Driven Semantic KV Retrieval for Long-Context LLMs
Abstract
Long-context inference with large language models is bottlenecked by the linear growth of the KV cache. Existing eviction-based methods reduce memory overhead by irreversibly discarding historical states, while retrieval-based methods preserve the full context but rely on retrieval units defined by tokens, fixed pages, or proxy signals including query-key similarity, clustering, and linguistic rules. Such externally defined units often misalign with the model's actual contextual dependencies during generation. We propose AsdaKV, an adaptive semantic drift-aware KV cache retrieval framework that defines semantic continuity directly from the model's inherent attention behavior. AsdaKV detects semantic drift by measuring the overlap between high-attention historical regions across decoding steps: stable overlap indicates that the active cache still covers the relevant contextual evidence, whereas a sharp overlap drop triggers cache refresh. Guided by this drift signal, AsdaKV organizes historical KV states into variable-length semantic windows, offloads full window KV states to CPU memory, and maintains only lightweight window anchors plus a fixed-budget active KV cache on GPU. During decoding, the current query retrieves relevant windows via anchor matching, while adaptive-stride drift detection and deferred page recall reduce redundant detection, retrieval, and KV reloading overhead. Extensive experiments across diverse tasks and model families demonstrate that under the same KV budget, AsdaKV achieves higher accuracy than state-of-the-art KV retrieval methods, maintains comparable decoding efficiency in long-input settings, and further outperforms existing retrieval systems for long-output generation.