GEAR: Bridging the Planner-Actor Gap via Gradient-Aligned Policy Extraction
Abstract
Hybrid model-based reinforcement learning (MBRL) integrates lookahead planning with actor-critic optimization for exceptional sample efficiency. However, this decoupled architecture suffers from a critical planner-actor mismatch. Existing policy alignment methods face a structural dilemma: forward KL triggers mode-covering objective conflicts, while reverse KL relies on brittle proxies that bottleneck expressivity. To address this challenge, we propose Gradient-Embedded Alignment Regularization (GEAR), a unified in-sample extraction framework. At its core, we derive a first-order geometric alignment metric from the continuous-time Hamilton-Jacobi-Bellman (HJB) equation to evaluate action evolution against the local value gradient. By embedding this metric as an absolute modulation weight, GEAR eliminates spurious imitation and transforms the inherently diffuse forward KL into a focused, mode-seeking objective. Evaluations on the challenging HumanoidBench suite demonstrate GEAR achieves significantly higher sample efficiency and asymptotic performance than state-of-the-art baselines. The project code is available at https://anonymous.4open.science/r/GEAR-5E80.