DELTA: Robustly Training Label-Conditional Diffusion Models with Weak Annotations
Abstract
While label-conditional diffusion models exhibit remarkable generative capabilities recently, their success heavily relies on massive, cleanly labeled datasets. In practice, categorical supervision is rarely perfect: it is often corrupted by noise, clouded by ambiguity, or partially missing. Training directly on such weak annotations severely degrades generation quality, and existing robust methods offer only fragmented solutions that depend on scarce auxiliary priors. To overcome these limitations, we propose DELTA, a unified framework that robustly trains diffusion models across all three weak annotation types without any external priors. By treating the unknown true label as a latent variable, DELTA optimizes a principled variational objective that jointly recovers the clean data distribution and infers the true label posteriors. To make this joint training computationally tractable, we further introduce a median-centered timestep sampling strategy that efficiently concentrates evaluation where the diffusion process is most informative. Extensive experiments demonstrate that DELTA produces high-fidelity, class-consistent samples across diverse weak supervision scenarios, outperforming specialized baselines while demanding strictly less prior knowledge.