DiffRatio: Training One-Step Diffusion Models Without Teacher Supervision
Wenlin Chen ⋅ Mingtian Zhang ⋅ Jiajun He ⋅ Zijing Ou ⋅ José Miguel Hernández-Lobato ⋅ Bernhard Schölkopf ⋅ David Barber
Abstract
Score-based distillation methods train one-step diffusion models in two stages: they first train a teacher score model, then distill it into a one-step student model. However, this distillation process can introduce bias from two sources: errors in the teacher score model and student score estimate. We propose DiffRatio, a new framework for training one-step diffusion models without teacher supervision. Instead of using a teacher score model to provide training targets, DiffRatio directly learns a log density ratio between the student and data distributions across diffusion time steps. As a pre-trained score model is only used to initialize the one-step generator and does not supervise training. This design simplifies the training pipeline, mitigates gradient estimation bias, and reduces the size of auxiliary networks. In addition, the learned density ratio can be used as a verifier, enabling a principled inference-time parallel scaling method that further improves sample quality without external rewards or extra sequential computation. DiffRatio achieves strong one-step generation results on CIFAR-10 and ImageNet ($64{\times}64$ and $512{\times}512$), outperforming most teacher-supervised distillation methods.
Chat is not available.
Successful Page Load