Iterative Gumbel Planning for Continuous Control
Shaohuai Liu ⋅ Weirui Ye ⋅ Yilun Du ⋅ Le Xie
Abstract
Search-based planning is a powerful policy improvement paradigm in reinforcement learning, with milestone successes on discrete-action domains through the AlphaZero and MuZero families. Extending it to continuous control, however, faces a structural mismatch between the search operator and the policy class: search at each state evaluates a finite candidate set and returns a discrete distribution, while the parametric Gaussian policy commonly used in continuous control is a continuous density. We prove that distilling the search output into a Gaussian induces a non-vanishing projection error that inflates policy variance and produces a return gap that persists even when the policy mean is locally optimal, and is not closed by additional search compute or data. Inspired by the iterative refinement of MPPI in continuous control, we propose \textbf{Iterative Gumbel Planning (IGP)}, which adapts this principle to a discrete policy class. IGP represents the policy as a per-dimension factorized categorical distribution aligned with the search support, eliminating the projection error and reducing the residual approximation to a bounded discretization error controllable by the grid resolution. Joint candidates are drawn by Gumbel Top-$K$ sampling without replacement, avoiding the $V^{|A|}$ combinatorial explosion of joint discretization, and search-refine iteration runs for multiple rounds at the same state. On DMControl and HumanoidBench, IGP achieves state-of-the-art performance, with improved sample efficiency and lower temporal action variance at evaluation.
Chat is not available.
Successful Page Load