Self-Cleaning Diffusion Models
Abstract
We present Self-Cleaning Diffusion Models, a principled and effective framework for training diffusion models from heterogeneous data. In many domains, high-quality data is scarce or expensive to collect, while low-quality and out-of-distribution (OOD) data is abundant. Leveraging these ubiquitous samples is critical for scaling generative models. Still, principled techniques remain elusive, leaving practitioners to rely on simple heuristics such as high-quality finetuning or explicit quality conditioning. Self-Cleaning Diffusion Models bridges this gap by decoupling data correction from prior learning. We first train a transport map between the abundant OOD and the limited in-distribution samples. We then train an unconditional generative prior on a mixture of the few in-distribution samples and noisy versions of the transported points. Crucially, this noise injection prevents learning errors in the transport map from propagating to the generative prior. Theoretically, we demonstrate that transporting OOD samples towards the empirical proxy reduces the distance to the underlying distribution under mild assumptions. Empirically, we achieve state-of-the-art results across a diverse range of settings, from mitigating synthetic corruptions to increasing diversity and transcending dataset quality.