Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating
Haozhuo Zhang ⋅ Qiang Zhang ⋅ Jian Tang ⋅ Mingzhe Ni ⋅ Michele Caprio ⋅ Angelo Cangelosi ⋅ Wei Pan
Abstract
Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy expands $N$ scenes into $S \times N$ concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero in later-stage streams if any previously placed object is displaced beyond a set threshold, thus maintaining object placement throughout the episode without extra reward terms. Collectively, these strategies achieve per-stage success rates exceeding 80\% on the LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes), with performance remaining stable even as the number of sequentially transported objects increases beyond two—unlike the sharp drop seen in all baselines. Additionally, we demonstrate that the RL teacher policy can be distilled into a Vision-Language-Action student via DAgger. When provided only with egocentric RGB or depth observations and natural language instructions, the student successfully replicates the teacher's long-horizon, multi-object behaviors, highlighting the potential for perception-driven deployment on real humanoid robots.
Chat is not available.
Successful Page Load