Disentangling the Good From the Bad: Quantization-Induced Flips Are Not Random
Abstract
Model quantization has advanced tremendously in recent years, producing models that require a fraction of the compute while maintaining baseline accuracy and creating the illusion of lossless compression. Beneath the surface, however, model behavior undergoes severe shifts that simple global accuracy fails to capture. Notably, quantized models experience a high rate of "flips": individual predictions that change from correct to incorrect (bad flips-CI) and vice versa (good flips-IC). Although overall accuracy remains stable, these changes degrade model reliability and introduce unforeseen safety risks. Current literature largely attributes these flips to random rounding noise, a perspective that ignores the specific precision constraints imposed by quantization with or without training. In this paper, we show that flips are driven by quantization's implicit regularization, low-pass filtering, and exacerbated spectral bias. By analyzing these effects via Singular Value Decomposition (SVD), we leverage the model's low-to-high rank internal disagreement to disentangle good flips from bad ones. Building on this insight, we introduce post-hoc approaches to selectively preserve the former while correcting the latter, and we validate the universality of our disentanglement through extensive experiments spanning Post-Training Quantization (PTQ) to Quantization-Aware Training (QAT) and across vision, language, and multi-modal models. With this work, we aim not only to restore the quantized model's hidden reliability degradation but also to elevate it beyond the accuracy ceiling of its full-precision counterpart.