How Diffusion Models Memorize
Abstract
Despite their success in image generation, diffusion models can memorize training data, raising serious privacy and copyright concerns. Although prior work has identified empirical factors associated with memorization, such as text guidance, prompt specificity, and duplicated training examples, the mechanism by which memorization is triggered and propagated during denoising remains unclear. In this paper, we provide a theoretical explanation of how memorization occurs in text-to-image diffusion models. We show that conditional overfitting causes the conditional posterior mean to collapse to the memorized training latent, while the unconditional posterior remains close to the zero-centered data mean at the first denoising step. As a result, classifier-free guidance overestimates the clean prediction and injects an amplified memorized signal into the reverse trajectory. We further show that this signal propagates across timesteps because each guided latent is fed back into the denoising network, allowing the conditional branch to indirectly shift the unconditional branch. To characterize this process, we introduce the posterior lag, the discrepancy between conditional and unconditional posterior means. For normal prompts, this lag decreases as the two posteriors synchronize during denoising. For memorized prompts, the conditional posterior commits early to a specific training latent, while the unconditional posterior catches up only later, producing a distinctive rise-and-fall lag pattern. Our analysis provides a mechanism-level explanation of memorization and clarifies why text guidance and early denoising steps play a central role in memorized generation.