Asymptotic Learning Curves for Conditional Diffusion Models with Random Features
Abstract
Conditional diffusion models generate diverse, high-quality, and novel samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, whereas recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term ``malign generalization.'' Furthermore, analyzing the training loss reveals that the more informative the condition, the more prone the model is to memorizing training samples. These theoretical findings are supported by experiments with U-Net architectures on realistic data.