Aligning Policy Encoders with a Shape Teacher
Angela Madelon Bernardy ⋅ Mohammad Mahdi Derakhshani ⋅ David Knigge ⋅ Riccardo Valperga ⋅ Artem Moskalev
Abstract
3D-aware representations can improve robotic manipulation by providing geometric information that is robust to changes in viewpoint, object position, and occlusion. However, using 3D inputs at deployment requires additional sensing and processing, making image-based policies more practical. We ask whether an image-based policy can acquire 3D shape information during training while retaining a purely image-based interface at deployment. We propose shape-teacher alignment, a training-time mechanism that uses a pretrained 3D shape representation as an alignment target for an image-based policy encoder. Specifically, we use Point-MAE features from individual objects as targets for corresponding image patches, transferring geometric structure to the policy without changing its deployment inputs or architecture. On LIBERO, shape-teacher alignment produces image representations that are more structured around object shape and corresponding parts, while substantially accelerating policy convergence and improving success rates. For pretrained policies, convergence is accelerated by 4.7$\times$ and 5.6$\times$ on the Object and Spatial suites, respectively, without requiring the shape teacher or depth input at deployment.
Chat is not available.
Successful Page Load