Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
Abstract
Recent advances in one-step text-to-image generation have enabled real-time synthesis with remarkable efficiency and quality. Previous reinforcement learning methods for one-step generators combines the image space reward optimization and the diffusion noisy space distribution matching. This paradigm brings challenges due to a mismatch between terminal reward optimization and the underlying generative dynamics. As a result, optimization tends to exploit stochastic degrees of freedom, often improving reward at the expense of image fidelity. To address this issue, we propose \textbf{Diff-Instruct with Diffused Reward} (\textsc{Didr}), a data-free trajectory-level alignment framework derived from Integral KL minimization. \textsc{Didr} propagates the RLHF-optimal reward-tilted clean-image distribution across all noise levels along the diffusion trajectory. We show that this objective admits the same minimizer as clean-image RLHF, while naturally inducing \textbf{the Diffused Reward Score} (DRS), which acts as a reward-driven correction to the reference score function. To make this practical, we further introduce \textbf{the Diffused Reward Proxy} (DRP), an efficient estimator of DRS based on differentiable short-step denoising. Extensive experiments demonstrate that \textsc{Didr} consistently Pareto-dominates existing one-step SDXL baselines. Moreover, when transferred to a 6B DiT backbone (\textsc{Z-Image}), \textsc{Didr} surpasses its 50-step teacher in preference alignment while requiring only a single generation step.