KV Cache Compression via Attention Output Distortion Minimization
Abstract
Efficient long-context inference requires compressing the KV cache under strict memory budgets without degrading model outputs. We formulate this problem as a constrained optimization: minimize the total attention-output distortion subject to a global average bit-width budget, where each token’s KV entry can be assigned a bit-width ranging from 0 (eviction), through mixed precision, to 16 (no compression). The central challenge is to estimate each token’s contribution to attention-output distortion. We address this via a first-order perturbation analysis of the attention mechanism, which yields a closed-form decomposition of the distortion. Specifically, the compression cost of each token factorizes into (i) a token-specific risk score that captures the sensitivity of the attention output to perturbations of that entry, and (ii) a precision-dependent quantization error determined by the chosen bit-width. Building on this factorization, we solve the global allocation problem via Lagrangian relaxation, reducing it to independent per-token decisions that can be efficiently computed. Experiments on LongBench and RULER with LLaMA-3.1-8B-Instruct and Qwen3-8B show that our method achieves comparable accuracy to uniform 4-bit quantization while using an average of 2 bits per entry (2× memory reduction). At an average of 1 bit, our method maintains substantially higher accuracy than both uniform quantization and token eviction baselines, which deteriorate sharply in this regime. The advantage concentrates on retrieval-sensitive subtasks, where dispersed evidence makes uniform allocation and hard eviction especially costly.