Debiased DPO for Diffusion Models
Abstract
Direct Preference Optimization (DPO) enables offline alignment of diffusion models from preference pairs without explicit reward modeling, but scaling DPO requires large amounts of expensive human preference data. In this paper, we consider a semi-supervised approach where only a small subset of preference pairs is annotated by humans, while the remaining (much larger) pool is unlabeled and annotated by inexpensive synthetic feedback e.g., vision-language models scoring image pairs or self-training from the current model. However, such synthetic supervision is not a replacement for human judgments: it can be systematically misaligned, so naively mixing human and synthetic preferences will yield misaligned inference. We introduce DeDPO, a debiased objective that integrates a doubly robust estimator from causal inference into DPO to correct synthetic-label misalignment while preserving the simplicity of DPO. Experiments show improved label efficiency and robustness to synthetic labeler quality, allowing DeDPO to outperform standard DPO under the same human-label budget and to approach fully human-supervised performance, despite using four times fewer human-labeled data.