SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework
Abstract
Building humanoid robots that perform generalizable whole-body loco-manipulation in the real world remains a fundamental challenge: existing approaches either rely on heavy task-specific reward engineering, rigidly replay reference motions that fail to generalize, or depend on costly teleoperation that limits scalability. While human videos capture diverse human behaviors, the motion priors inferred from them are inherently imperfect, suffering from occlusion, contact artifacts, and retargeting errors that render them unsuitable for direct policy learning. To this end, we present SUGAR, a data-driven framework that converts diverse human videos into deployable humanoid loco-manipulation skills, without any task-specific reward engineering or reference-motion conditioning at inference. SUGAR proceeds in three coupled stages: First, a fully automated pipeline extracts scalable kinematic interaction priors including human-object motion trajectories and contact labels from diverse human videos. Second, a privileged physics-based refiner utilizes a unified mimic-style reward and a progressive state pool to transform imperfect kinematic interaction priors into physically feasible, high-fidelity skills. Third, the refined skills are distilled into a hierarchical policy that comprises a task-guided planner and a command tracker for autonomous task execution. We evaluate our method on six representative loco-manipulation tasks in both simulation and real-world humanoid hardware. SUGAR substantially outperforms reference-tracking baselines, and its performance scales clearly with the amount of human video data. It also achieves zero-shot real-world transfer with reliable closed-loop execution, autonomous failure recovery and stable long-horizon performance under external perturbations. Project Page: https://sugar-humanoid.github.io/