ADKV: A Low-Overhead Adaptive Delta Quantization for KV Cache in LLM Inference
Honghao Jia ⋅ Zhenxing Li ⋅ Jialiang Guo ⋅ Jiacheng Gan
Abstract
The KV cache is a major memory bottleneck in LLM inference, especially under large batch sizes and long contexts. While quantization is a promising solution, existing static group-wise quantization methods are constrained by a trade-off between metadata overhead and accuracy: increasing the group size reduces per-group metadata (e.g., zero-points, scaling factors, and outlier indices), but exacerbates quantization errors under the sinusoidal-like temporal drift of post-RoPE outlier channels; conversely, using smaller groups improves accuracy but incurs substantial metadata overhead. In this paper, we propose ADKV, a low-overhead adaptive delta quantization framework for KV cache in LLM inference, which leverages the temporal structure of the KV cache. ADKV tracks per-channel zero-points and scaling factors via online EMA, adapting to the sinusoidal-like temporal drift of post-RoPE outlier channels, thereby eliminating the rigidity of static group-wise quantization. This adaptive mechanism further enables delta coding to embed metadata directly into the quantized sequence, decoupling metadata overhead from context length. To minimize the MSE of attention outputs, we design an SGD-based calibration algorithm that models ADKV as an RNN. Experiments on generation and long-context benchmarks show that, while matching the accuracy of static group-wise quantization, ADKV reduces the metadata memory footprint by $\sim 12.8\times$. A custom fused ADKV-GEMV CUDA kernel optimized for memory-bound autoregressive generation achieves $\sim 1.96\times$ throughput over the FP16 baseline.
Chat is not available.
Successful Page Load