ResKV: Residual-based Channel-wise Unstructured Pruning for KV Cache Compression
Abstract
In modern LLM applications, the demand for extended context lengths continues to surge, making KV cache a major performance bottleneck in long-context inference. Channel-wise pruning has emerged as a prevalent KV cache compression paradigm; however, existing solutions overlook inherent similarities among KV vectors, resulting in limited pruning ratios and noticeable performance degradation. To tackle this issue, we present ResKV, a novel residual-driven channel-wise pruning framework that fully exploits KV vector similarity for effective compression. ResKV first clusters highly similar KV vectors via cosine similarity metrics, then calculates residual vectors by subtracting each cluster’s centroid from the original features. Subsequent channel-wise unstructured pruning and per-token quantization are exclusively applied to these residual components. During inference, preserved residuals are combined with their corresponding centroids to accurately recover dense KV representations. Benefiting from the low magnitude and sparse, near-zero properties of centroid residuals, both pruning and quantization incur only negligible approximation errors. Experimental results show that ResKV reduces peak memory consumption by 45.6% and improves inference throughput by 4.27× against dense FlashAttention inference, while sustaining competitive model performance. Under identical sparsity constraints, ResKV outperforms state-of-the-art channel-wise KV pruning methods by achieving up to 2.26 points higher accuracy and 2.63× faster end-to-end throughput.