Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
Jérémie Klinger ⋅ Raphaël Urfin ⋅ Giulio Biroli ⋅ Marylou Gabrié
Abstract
Score-based generative models generate new samples by integrating a time-dependent velocity field that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weighting schedule $w(t)$. Along the backward dynamics, the trajectory commits to one mode of the target within a narrow window, the *speciation time*. This holds for an exact drift, leaving open how structure emerges during the training dynamics. In this work, we show that $w(t)$ sets the rate at which each feature of a multimodal target -- the modes and their relative weights -- is acquired, and that these rates are governed by the signal-to-noise ratio $\Lambda(t)$ at which the network is trained. Crucially, at high $\Lambda(t)$ all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are never learned. It is only near the speciation time, where $\Lambda(t)$ becomes order 1, that the features are acquired on distinct timescales: the weights become learnable, jointly with the directions, and the directions themselves are learned at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is governed by how much of the weighting effectively sits near the speciation time. Building on the exact high dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures, we show that these results actually hold in more complex settings, including image or human genome haplotypes generation.
Chat is not available.
Successful Page Load