CMPQ: Compensatory Quantization via Input-Aware Hessian Damping
Abstract
With the recent advancement of large language models (LLMs) scaling, post-training weight quantization (PTQ) has become one of the standard tools for making models deployable on commodity hardware with minimal performance degradation. However, existing quantization solvers such as GPTQ calibrate each weight matrix against a corrupted signal: the calibration input observed at any deep layer is the output of the cumulative quantized chain, not a clean full-precision reference. We propose \textbf{CMPQ (Compensatory Quantization)}, which reframes the corruption as evidence about \emph{which input directions have become unreliable} and tells the solver to stop spending its compensation budget on them. CMPQ measures the divergence between the quantized and full-precision input and uses it to reshape the curvature that GPTQ already relies on, redirecting compensation onto weights attached to clean input subspaces and cutting the propagation path of upstream noise. On the Qwen3 family at three sizes (4B / 8B / 14B), CMPQ matches GPTQ-family baselines at 3-bit and substantially outperforms them at 2-bit, where standard solvers collapse to near-random accuracy.