ClusQuant: Mitigating Outliers with Clustering-Based Representations for Low-Precision LRMs
Xingyu Liu ⋅ Xiangyang Yin ⋅ Tianhua Xia ⋅ Haiyu Wang ⋅ Sai Qian Zhang
Abstract
Large reasoning models (LRMs) improve task performance by generating longer intermediate reasoning traces, but this substantially increases inference costs and makes low-precision deployment increasingly important. However, aggressive quantization often performs poorly on reasoning tasks. Beyond error accumulation across autoregressive generation, we observe that activations exhibit significantly higher variability during the early decoding steps of thinking, making conventional calibration and uniform quantization less effective. To address these reasoning-specific quantization challenges, we propose \textbf{ClusQuant}, a clustering-based LUT quantization framework for low-precision large reasoning models. ClusQuant strengthens outlier mitigation through rotation and scaling transformations, uses clustering-based representations to better fit non-uniform activation distributions, and replaces fixed codebook sizes with an adaptive clustering strategy that allocates codebook capacity according to layer-wise quantization difficulty. Guided by our observation of reasoning-time activation behavior, ClusQuant further improves calibration sample selection to better match actual decoding statistics. Combined with our customized LUT CUDA kernel, ClusQuant delivers substantial accuracy improvements on reasoning benchmarks while achieving $2.85\times$ speedup.
Chat is not available.
Successful Page Load