A First-Order Theory of Decision Flips Under Post-Training Quantization: What It Predicts, and Where It Breaks
Abstract
Post-training quantization (PTQ) is applied to language models with no safety objective in mind, yet it measurably shifts individual decision boundaries, including, in our case study, a model's refusal decision. We ask whether a first-order perturbation theory, one that can be estimated from forward passes alone before any quantization happens, can predict which decisions will flip. Treating quantization as a bounded weight perturbation, we calibrate a variance model against real per-format quantization error measured on six checkpoints and find that it holds within a factor of two. The matching zero-crossing flip criterion, however, under-predicts the observed flip rates by 3.5 to 454 times. We trace the gap to structured anisotropy in real quantization error, a pattern the isotropic-perturbation assumption cannot capture, and show that a much simpler quantity, the pre-quantization margin itself, needs no perturbation model at all and still beats the theoretically motivated predictor by a wide margin (AUC 0.95 versus 0.69). We present this as a case study in where current first-order theory for PTQ works, where it breaks down, and why.