Analyzing Learning-Dynamics Across Loss and Prediction Choices for Diffusion Models
Dongwoo Kim
Abstract
Denoising generative models predict clean data ($x$), noise ($\epsilon$), or velocity ($v$), and are trained with a squared loss that can be computed in any of these three spaces. Prior work treats these choices as interchangeable up to loss reweighting, but practitioners observe sharp qualitative differences that remain unexplained. We systematically decouple the \emph{prediction space}, what the network outputs, from the \emph{loss space}, where supervision is applied, creating a $3{\times}3$ design matrix. We show that loss-space changes only reweight how much each noise level contributes to training, while prediction-space changes alter what the network is asked to learn: two fundamentally different mechanisms. This decoupling exposes two regimes. Without a representational bottleneck, the loss choice primarily controls denoising quality (MSE) while the prediction choice primarily controls sample quality (FID). The reason is that ODE sampling must convert the network's output back into a clean image, and this conversion amplifies errors differently across prediction types, turning sub-one-percent gaps in training error into 22-fold gaps in FID that are invisible to standard training diagnostics. Under a representational bottleneck, the prediction choice instead dominates both metrics: recovering noise or velocity requires reconstructing high-dimensional information that the bottleneck has destroyed, while recovering the clean image remains feasible because real images lie on a low-dimensional manifold. We validate these findings on a U-Net and the full 130M-parameter JiT architecture across CIFAR-10 and Imagenette, and give architecture-dependent recommendations for both regimes.
Chat is not available.
Successful Page Load