Provably Reliable Classifier Guidance via Cross-Entropy Control
Sharan Sahu ⋅ Arisina Banerjee ⋅ Yuchen Wu
Abstract
Classifier-guided diffusion models generate conditional samples by augmenting the reverse-time score with the gradient of the log-probability predicted by a probabilistic classifier. In practice, this classifier is usually obtained by minimizing an empirical loss function. While existing statistical theory guarantees good generalization performance when the sample size is sufficiently large, it remains unclear whether such training yields an effective guidance mechanism. We study this question in the context of cross-entropy loss, which is widely used for classifier training. Under mild smoothness assumptions on the classifiers, we show that controlling the cross-entropy at each diffusion model step is sufficient to control the corresponding guidance error. In particular, we show that probabilistic classifiers achieving conditional KL divergence $\varepsilon^2$ induce guidance vectors with mean squared error $\widetilde O(d \varepsilon )$, up to constant and logarithmic factors. We demonstrate that the proposed smoothness condition is necessary, by constructing a sequence of non-smooth classifiers that achieve small conditional KL while inducing an exploding guidance error. We also show that the derived upper bound $\widetilde O(d \varepsilon )$ is optimal up to poly-logarithmic factors. Our result yields an upper bound on the sampling error of classifier-guided diffusion models and bears resemblance to a reverse log-Sobolev--type inequality. To the best of our knowledge, this is the first result that quantitatively links classifier training to guidance alignment in diffusion models, providing both a theoretical understanding of when such models succeed, along with principled guidelines for selecting classifiers that induce effective guidance. In particular, our findings suggest that, when selecting classifiers for guidance, one should prioritize not only accuracy but also smoothness, highlighting the advantages of smoothness-inducing training methods for diffusion guidance.
Chat is not available.
Successful Page Load