Look-ahead Variational Flow for Generative Online Reinforcement Learning
Abstract
Generative policies have emerged as a compelling paradigm in reinforcement learning (RL) as they can represent complex and multimodal action distributions. However, this high expressivity leaves them particularly vulnerable to the inherent target distribution non-stationarity in RL, rendering the generative process severely volatile. To address this issue, we propose \textbf{L}ook-ah\textbf{E}ad v\textbf{A}riational \textbf{F}low (LEAF), a generative policy optimization framework that stabilizes both the target and the generative process. First, we replace the action-value-guided target with a look-ahead target induced by predictive dynamics, which we theoretically prove to exhibit a tighter target drift bound. Second, we formulate policy improvement as a proximal variational free-energy optimization, which enforces a distributional trust region to prevent overfitting to transient target shifts. Finally, we prove that this optimization induces an optimal transport flow, which we instantiate as a two-phase coarse-to-fine generative process to achieve broad multimodal coverage and pinpoint accuracy. Extensive experiments on the DeepMind Control Suite and Humanoid Bench demonstrate that LEAF consistently outperforms strong baselines in both sample efficiency and asymptotic performance.