Mitigating Compounding Errors in Online Reinforcement Learning via Optimal Transport Regularized Flow Matching
Abstract
Model-based reinforcement learning (MBRL) promises high sample efficiency but is often crippled by compounding model errors: small prediction inaccuracies accumulate rapidly over multi-step rollouts, leading to biased policy optimization. We propose FTA, a trajectory-level generative framework using conditional flow matching (CFM) instead of step-wise dynamics. To avoid curved, energy-inefficient paths, we introduce an optimal transport regularizer grounded in the Benamou–Brenier formula. This regularizer minimizes kinetic energy while enforcing exact state displacement, yielding straight, low-energy trajectories with provable linear error propagation. Since a static generator fails in online RL due to (i) shifting data distributions and (ii) mismatch between training (historical) and desired (high-return) conditionals, we add periodic retraining from the replay buffer and a value-guided ODE sampler that biases generation toward high returns via the current Q-function. Theoretically, the regularizer drives the flow toward the Wasserstein-2 optimal transport map, and value-guided sampling reduces asymptotic Q-estimate bias. On MuJoCo benchmarks, FTA consistently outperforms strong MBRL and model-free baselines in sample efficiency and final performance, and generalizes across tasks. By mitigating compounding errors at the trajectory level and adapting to online shifts, FTA offers a robust, data-efficient alternative to traditional dynamics models.