Same Score, Different Model: Two Regimes of Quantization Damage Below and Above the Accuracy Plateau
Abstract
Quantized models are usually validated by checking that a benchmark score did not drop. We show this test is insensitive by construction, and that what it misses differs qualitatively across the precision range. Evaluating seven GGUF quantizations of a 30B-parameter Mixture-of-Experts model on a fixed 5,000-question STEM subset of MMLU-Pro (3,658 after excluding responses truncated by the generation budget), we find accuracy plateaus at ~3.9 bits per weight, with Q80 and BF16 scoring identically. Yet agreement with BF16, the fraction of questions answered the same way, rises monotonically at every step over the same range, from 0.927 to 0.969. The disagreements behind that gap fall into two regimes: below the plateau they are asymmetric and quantization loses roughly twice as often as it wins (-61 net at UD-IQ1M), while above it the asymmetry is no longer detectable (42/42 at Q80), so the model becomes a different model of indistinguishable quality rather than a degraded one. Accuracy is blunt because only 12.2% of questions are contested at all. We further show that the BF16 answer-token margin predicts which questions flip, and that same-model sampling variance (16.4%) exceeds the entire Q80-BF16 gap (3.1%). All per-question outputs are released.