DICEQuant: Distortion-Compensated Rounding with Dual-Ended Shrinkage for LLM Quantization
Yuan Cheng ⋅ Xing Hu ⋅ Zukang Xu ⋅ Hui Wang ⋅ Xiaomeng Han ⋅ Dawei Yang
Abstract
Post-Training Quantization (PTQ) has become a prerequisite for efficient LLM deployment; however, current rotation-based methods are approaching a saturation point because they prioritize heuristic approximations over theoretical rigor. By treating quantization as an opaque black box, existing paradigms typically depend on indirect gradient estimators for activations while restricting weight optimization to local reconstruction objectives. Consequently, this structural limitation induces fundamental deficiencies: specifically, optimization instability and a "twofold blindness" toward accumulated upstream distortion and final task loss sensitivity. To overcome these barriers, we introduce $\textbf{DICEQuant}$, a unified framework established upon rigorous error modeling. First, we propose the $\textbf{CURE (Coupled Underflow-Rounding Error) surrogate}$, which facilitates exact and stable gradient computation for activation reshaping, thereby eliminating the variance inherent in heuristic estimators. Simultaneously, for weight optimization, we present $\textbf{Distortion-Compensated Rounding (DCR)}$. This mechanism derives a Shifted Optimal Center to neutralize input noise and utilizes Hessian-weighted shaping to align rounding decisions with the global functional objective. Empirical results demonstrate that DICEQuant significantly surpasses state-of-the-art baselines, successfully bridging the accuracy gap between low-bit compression and full-precision intelligence.
Chat is not available.
Successful Page Load