AdaKVQ: Adaptive Mixed-Precision KV Cache Quantization For Efficient Reasoning Models
Abstract
Large Reasoning Models (LRMs) rely on extensive Chain-of-Thought generation but suffer from memory bottlenecks due to linear Key-Value (KV) cache growth. While low-bit quantization offers a potential solution, we identify that current static quantization strategies suffer from severe degradation on complex reasoning tasks at extreme bit-widths. Notably, this accuracy drop is accompanied by increased generation length, which ironically limits the intended efficiency gains. We attribute this performance collapse to the fact that static quantization methods are insufficiently flexible to meed dynamic demands of precise long-range retrieval. Utilizing Average Attention Distance (AAD) to quantify retrieval patterns, we find that the attention patterns in shallow layers exhibit strong locality, whereas that in deep layers demonstrate intensive long-range dependencies. Furthermore, deep layers have fluctuating AAD values across token positions. Based on above findings, we propose AdaKVQ, an adaptive mixed-precision quantization framework that modulates the bit-width allocation guided by the dynamic demands of long-range context retrieval in each layer. Extensive experiments on complex reasoning benchmarks with Llama-3.1 and Qwen2.5 demonstrate that AdaKVQ consistently outperforms state-of-the-art baselines, achieving performance comparable to full-precision models.