UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
Abstract
Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that learns a unified physical language for human-to-humanoid manipulation transfer. Grounded in the philosophy that heterogeneous kinematics share consistent visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch integrates these purified modalities into a shared discrete latent space of cross-embodiment physical intents. We validate UniT across two paradigms: (1) Policy Learning (VLA-UniT): By predicting these unified tokens, VLA-UniT achieves state-of-the-art performance with high data efficiency on the RoboCasa GR1 benchmark. Leveraging diverse human data further improves out-of-distribution (OOD) generalization in simulation and real-world deployment, and enables zero-shot task transfer in the real world. (2) World Modeling (WM-UniT): By aligning cross-embodiment dynamics via unified tokens as conditions, it supports direct human-to-humanoid action-conditioned generation, translating human knowledge into enhanced action controllability for humanoid video generation. Ultimately, by inducing a more aligned cross-embodiment representation (empirically supported by t-SNE visualizations revealing improved alignment of human and humanoid features), UniT offers a scalable path to distill human priors into humanoid manipulation capabilities.