InQuant: In-Place Mixed-Precision KV Cache Quantization via Saliency-Aware Neighbor-Slot Reuse
Zihan Chang ⋅ Shuibing He ⋅ Bo Zhou ⋅ Ping Chen ⋅ Siling Yang
Abstract
Key-value (KV) caches are essential for efficient Large Language Model (LLM) inference, but their memory footprint grows linearly with context length and batch size. Low-bit KV-cache quantization reduces this footprint, yet uniform quantization is vulnerable to high-magnitude outlier channels, while outlier-aware mixed-precision methods often introduce extra buffers, indexing, or channel permutation that weakens their system-level benefit. This paper presents InQuant, an in-place mixed-precision KV-cache quantization method that preserves salient outlier channels without changing the physical 4-bit packed layout. The key observation is that selected low-saliency channels near outliers can tolerate small approximation error; InQuant therefore reuses their storage slots to hold the extra 4-bit nibble needed by 8-bit outlier values. To make this layout practical, InQuant combines sampling-based channel saliency estimation, saliency-aware neighbor-slot reuse, stride-based handling for adjacent outlier groups, and descriptor-guided marker recovery during dequantization. Across six LLMs and long-context workloads, InQuant reaches a fixed 4.0$\times$ physical KV-cache compression ratio, preserves accuracy close to strong mixed-precision baselines, and reduces quantization/dequantization latency by 2.2$\times$--3.4$\times$ compared with representative outlier-aware methods.
Chat is not available.
Successful Page Load