Why Classifier-Free Guidance Is Fragile Under Quantization: A Mechanistic Study
Ahmad Imran
Abstract
Classifier-free guidance (CFG) combines two denoiser predictions through their difference, but diffusion post-training quantization calibrates each branch's error independently of this coupling. We show this mismatch is not benign. If $\delta_u$ and $\delta_c$ are the quantization errors on the unconditional and conditional branches, the guided error is exactly $\delta_u+w(\delta_c-\delta_u)$: error the branches share passes through once, while error they disagree on is amplified by $w$. On a W4A4 SVDQuant baseline, three text-to-image models place a disproportionate share of their quantization error in precisely this amplified channel, giving $4.0$--$4.9\times$ guided-error amplification at default guidance scales, even though the branches' full-precision outputs agree almost perfectly. Controls on unrelated prompt/latent pairs and a $1/h$ scaling law rule out that this is an artifact of comparing two similar outputs: pollution, our diagnostic for it, is $17$--$82\times$ higher on real CFG pairs than on the controls, and grows as the true branch difference shrinks. A matched-energy noise-injection experiment on all three models confirms the effect is causal for image distortion: noise placed only in the disagreement channel does $3.4$--$6.7\times$ the LPIPS damage of the same-energy noise placed in the shared channel. Activation precision helps, but only up to a model-dependent ceiling. The result is a measurement target (branch-error alignment) that per-branch reconstruction objectives cannot see.
Chat is not available.
Successful Page Load