DrPO: Drifting Preference Optimization for One-Step Generative Models
Abstract
While one-step text-to-image generators, a fast and deployment-friendly class of visual generative models, can produce samples with a single forward pass, existing preference-finetuning methods fail to simultaneously achieve efficient adaptation, reward-model agnosticism, and diverse generation. In this work, we propose Drifting Preference Optimization (DrPO), an online preference-finetuning method for one-step generators. Inspired by the recent work of Drifting Models, DrPO first uses the target reward model to construct an on-policy dipole reward model from positive and negative samples. This dipole model induces a preference drifting field in latent feature space, whose gradients are then used to optimize the generator without backpropagating through the target reward model. Empirically, we evaluate DrPO across multiple reward models and SD-Turbo and SDXL-Turbo one-step generators, including HPSv3 and GenEval as representative alignment benchmarks, and validate that DrPO can efficiently and effectively finetune one-step generative models. Preliminary experiments further suggest that this sample-based gradient synthesis can be extended to offline preference finetuning.