Learning Dynamics of Convergence and Generalisation in Diffusion Models
Abstract
Defining generalisation is difficult in generative models, where held-out loss does not necessarily capture the quality and diversity of generated samples. An influential recent proposal characterizes generalisation through convergence: models trained independently on disjoint subsets of the same dataset produce similar samples from the same latent variable. While convergence has been studied empirically in deep generative models and theoretically as a function of sample complexity, its dynamics and relation to other measures of generalisation remain poorly understood. Here, we use random matrix theory to study convergence and generalisation during training in a linear diffusion model. We derive analytical predictions for sample reproducibility and latent-feature recovery, showing that these notions can emerge at distinct training times: convergence can peak before test-based criteria are optimized, while latent structure follows a separate dynamical transition. Their ordering depends on data structure and weight initialization, and can exhibit biased generalisation, with optimal generalisation occurring after overfitting begins.