Information-Theoretic Scaling Laws for Generative Diffusion Models
Abstract
In the infinite-data limit, diffusion models trained on different samples from the same data distribution should converge to the same learned distribution. However, in practice, a model trained on finite data depends on the particular training dataset; this dependence is reflected in properties of the model's outputs, such as memorization and sample quality. To understand this dependence through an information-theoretic lens, we introduce model information, a computationally tractable measure of how much samples from the trained model depend on the particular training dataset. We explore its scaling behavior as model size and training dataset size vary, revealing that it tracks distinct memorization and generalization regimes. Specifically, we uncover two qualitatively distinct scaling regimes: (1) a nonlinear model information regime marking the transition from pure memorization to imperfect generalization, and (2) a generalization regime characterized by log-linear scaling of the model information and progressively improved generalization and sample quality. Our results demonstrate information-theoretic principles for memorization and sample quality scaling laws in diffusion models.