Lift, See, Act: Hierarchical Robot Policy Pretraining with 3D Foundation Models
Yiyuan Ge ⋅ Changxing Ding ⋅ Ziyu Hao ⋅ Zijie Zheng ⋅ Xiangmin Xu
Abstract
Robot policy pretraining based on human videos is crucial for improving the policy’s generalization ability. One main challenge for this task is the lack of explicit action-relevant representations in such unlabeled data. Recent works tend to estimate 3D hand motion trajectories from videos using 3D foundation models (3D FMs). However, they keep solely the sparse trajectories as supervision, ignoring the rich, fine-grained information produced during the trajectory extraction process. To address this issue, we propose $\textbf{LSA}$, a hierarchical robot policy pretraining framework that includes three phases: $\textbf{L}$ift, $\textbf{S}$ee, and $\textbf{A}$ct. Specifically, the Lift phase elevates 2D video observations into dense 3D representations by adaptively aligning features in the policy's shallow layers with the intermediate representations extracted from multiple 3D FMs. Building upon these lifted representations, the See phase equips the policy with explicit geometric and interactive awareness through dual guidance. It introduces depth-map reconstruction to help the model comprehend 3D spatial layouts and utilizes hand region cues to explicitly supervise the encoder’s attention toward task-relevant interaction hotspots. Empowered by the dense 3D representation learning and precise spatial guidance, our policy achieves robust and accurate robotic manipulation in the Act phase. Extensive experiments on diverse simulation and real-world tasks demonstrate that LSA significantly outperforms current state-of-the-art approaches. The code of this work will be released soon.
Chat is not available.
Successful Page Load