LAPrune: Logits-Aligned Scoring Proxy for KV Pruning via Vector Quantization
Mingyang Yu ⋅ Rong-Cheng Tu ⋅ Yifu Ding ⋅ Hanqing Zhao ⋅ Yongcheng Jing ⋅ Xiao Luo ⋅ Dacheng Tao
Abstract
Long-context inference in Large Language Models (LLMs) suffers from high latency because attention computation and memory traffic scale linearly during decoding. KV cache pruning reduces this cost by retaining only a subset of cached tokens, but its effectiveness depends on whether the pruning proxy can recover the tokens with the largest exact attention logits under the current query. Existing proxies suffer from ranking misalignment: static heuristics rely on historical statistics and thus ignore the current query, while hashing-based retrieval proxies rank candidates in a discrete metric space that does not coincide with attention's inner product. To avoid these problems, we propose LAPrune, a KV pruning framework that uses additive vector quantization to build an efficient attention-logit-aligned proxy. LAPrune represents each cached key as a sum of learned codewords and decomposes the query-key dot product into lookup-and-add operations over query-codeword inner products. This replaces high-dimensional dot products with lightweight scalar lookups while preserving fidelity to the original attention ranking. LAPrune requires no base-model finetuning and can be integrated into long-context decoding with minimal modification. Across long-context evaluations, LAPrune improves exact-logit Top-$k$ recovery and perplexity over hashing-based pruning proxies, achieving up to $2.73\times$ decoding speedup over FlashAttention-2 and $4.83\times$ end-to-end speedup over the vanilla baseline at 128K context. Our code is available at [Anonymous/LAPrune](https://anonymous.4open.science/r/LAPrune-VQ).
Chat is not available.
Successful Page Load