Likelihood-Free Generative Policy Optimization
Abstract
Diffusion and flow policies have become promising tools for robotic and sequential decision-making systems. Many stable and effective policy optimization methods require evaluating the likelihood or likelihood ratio of an action under the current policy. However, these quantities are generally intractable for diffusion and flow models, which prevents widely used policy update algorithms from being applied to generative policies. We propose Likelihood-Free Generative Policy Optimization (LFGPO), a framework that avoids exact likelihood evaluation by separating policy improvement from generative model training. In the first stage, LFGPO uses a lightweight neural network to represent the likelihood ratio that carries the policy improvement signal. In the second stage, it distills the ratio into diffusion or flow policies through a weighted score- or velocity-matching objective. Our framework incorporates a large family of Reinforcement Learning (RL) algorithms, including PPO and GRPO, to be applied to diffusion and flow policy optimization. We show that the solution to the weighted score matching objective solves a stochastic optimal control problem. Across MuJoCo continuous-control benchmarks, LFGPO delivers strong and stable improvements over competitive diffusion- and flow-policy baselines, showing a practical route to RL fine-tuning for generative policies.