Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow Models
Andreas Bergmeister ⋅ Stefanie Jegelka ⋅ Nikolas Nüsken ⋅ Carles Domingo i Enrich ⋅ Jakiw Pidstrigach
Abstract
Diffusion and flow-matching models scale because pretraining simply regresses against a closed-form target built by analytically noising clean data samples. In many settings, only a reward function is available, scoring how desirable a generation is, so training must proceed online from a pretrained model. Existing methods either rely on costly SDE rollouts, sometimes with reward gradients, or adopt heuristic reward-dependent variants of the pretraining objective. Under a stochastic optimal control formulation of KL-regularized reward maximization, the optimal generative process tilts only the clean-endpoint distribution and leaves the conditional noising law unchanged. Combining this path-space characterization with the adjoint-matching optimality condition and a REINFORCE estimator to avoid reward gradients, we derive Reinforce Adjoint Matching (RAM). At each step, we draw a clean endpoint from the current model with any off-the-shelf sampler, evaluate its reward, noise it to several training states, and regress against a closed-form, reward-weighted target. Like the pretraining objective, RAM is simple and scales. On Stable Diffusion 3.5M, RAM achieves the highest reward on composability, text rendering, and human preference. It is much more training efficient, reaching Flow-GRPO's peak GenEval accuracy in~$50\times$ fewer training steps.
Chat is not available.
Successful Page Load