Amortizing Generative Guidance for Model-Based Reinforcement Learning
Abstract
Planner-guided model-based reinforcement learning (MBRL) has emerged as a powerful paradigm for continuous control, combining learned world models with online planning to achieve strong performance and sample efficiency. However, existing methods typically distill planner-improved actions into a unimodal Gaussian policy and reuse it as a proposal prior for subsequent planning, overlooking the fact that Model Predictive Path Integral (MPPI) planners can induce multimodal supervision over high-value actions. Compressing such supervision into a unimodal distribution can lead to mode averaging, limited proposal coverage, and unstable bootstrapped policy learning. To address this issue, we propose GeMP, a \textbf{Ge}nerative \textbf{M}ultimodal \textbf{P}lanning framework for MBRL. We theoretically characterize the multimodality of planner-induced supervision, showing that finite-sample planning can provide a distribution of high-value actions rather than a single deterministic target. To capture this distribution without compromising training stability, GeMP jointly learns a multimodal flow policy through value-weighted supervision and retains an auxiliary Gaussian policy for stablizing off-policy actor-critic learning. We further employ a heterogeneous proposal generator that combines Gaussian, flow-based, and random action-sequence candidates to improve proposal coverage for MPPI refinement. Experiments on MyoSuite and DeepMind Control Suite show that GeMP improves control performance and training stability over competitive baselines.