KVFocus: A Perturbation-Theoretic Token-Risk Score for Selective KV Cache Reuse in RAG
Shizhuo Zhang ⋅ Nuowen Kan ⋅ Chenglin Li ⋅ Rui-Xiao Zhang ⋅ Yifan Zhao ⋅ Yuanpeng He ⋅ Huixin Zhang ⋅ Fan He ⋅ Hang Xu ⋅ Hao Zhang ⋅ Wenrui Dai ⋅ Junni Zou ⋅ Hongkai Xiong
Abstract
In retrieval-augmented generation (RAG) scenarios, the prefill delay of the large language model (LLM) long inputs is the dominant component of time-to-first-token (TTFT). Reusing KV caches across different LLM inputs reduces this cost, but the resulting cross-context mismatch on reused segments has a significant degradation on answer quality via propagating through prefill attention into decode-time logits. Existing selective-recomputation methods mitigate this issue by recomputing tokens chosen by empirical importance heuristics, which overlook the propagation of cache mismatches into answer-side errors, thereby failing to effectively strike the quality-TTFT tradeoff. To address this issue, we propose a perturbation-theoretic KV cache reuse framework, KVFocus, which selects reused KV caches through a perturbation-theoretic token-risk score during the prefill process. Specifically, we first develop a first-order analysis that traces reuse error in three stages: source mismatch on the reused segment, suffix contamination during prefill, and decode-time logit stability. This analysis reveals that the leading-order answer-side error decomposes into a product of source-side V-drift and downstream suffix-to-segment attention concentration. Guided by this multiplicative structure, we define a token-risk score that operationalizes the bound during the inference. A plug-and-play selector for KV cache reuse in existing RAG-based LLM serving is then designed by recomputing the top-$r\%$ tokens directly suppresses the predicted leading-order answer-side error, covering both sides of the pathway that prior one-sided heuristics miss. Empirical results across three 7-8B instruction-tuned LLMs (Mistral, Qwen2.5, Qwen3) and four multi-hop QA benchmarks demonstrate that the proposed KVFocus achieves a superior trade-off between the ansewer quality and the TTFT in comparison to state-of-the-art baselines, with up to $3\times$ TTFT speedup over full prefill at matched answer quality.
Chat is not available.
Successful Page Load