Self-Adjoint Flow Policy Optimization
Abstract
Generative policies offer a promising alternative to Gaussian policies in reinforcement learning (RL) due to their ability to represent expressive and multimodal action distributions. However, integrating such policies into likelihood-based policy optimization remains challenging, since accurate and differential action likelihoods are intractable for standard generative policies. In this paper, we propose \emph{Self-Adjoint Flow Policy Optimization} (\textsc{SAFlow}), a likelihood-tractable generative policy with self-adjoint structure. \textsc{SAFlow} parameterizes the policy as an invertible flow in a doubled action space and uses a Verlet-style self-adjoint composition for generation. This structure makes the inverse map obtainable by stepsize reversal, aligning forward sampling and backward likelihood evaluation under the same time-symmetric numerical rule. It also provides a second-order approximation accuracy to the learned continuous flow, improving the numerical consistency of likelihood and policy-ratio computation. We instantiate \textsc{SAFlow} within an on-policy optimization framework and evaluate it on eight IsaacLab continuous-control tasks. \textsc{SAFlow} consistently outperforms Gaussian-policy baselines and recent generative-policy methods, highlighting the value of self-adjoint numerical structure for likelihood-based generative policy optimization.