RoboExo: Structure-Guided Wrist-to-Exocentric Video Generation for Scalable Robot Learning
Abstract
Scaling Vision-Language-Action (VLA) policy training requires diverse robot demonstration data, yet collecting such data remains expensive and labor-intensive. The community has explored ways to accelerate robot data collection: visual augmentation methods synthesize novel samples by editing visual elements in existing demonstrations, while systems such as UMI leverage wrist-mounted cameras to flexibly collect in-the-wild manipulation videos. However, the mismatched observation setups of these two directions leave them largely disconnected, limiting their synergistic potential for scalable robot data generation. To bridge this gap, we propose RoboExo, a generative framework for controllable wrist-to-exocentric conversion of robotic demonstrations. We introduce a geometry-motion factorization strategy that reconstructs canonical object-robot geometry and propagates motion states from noisy wrist-view videos, yielding reliable structural priors for cross-view synthesis. These priors guide a conditional video diffusion model to generate target-view exocentric videos that are spatially aligned, temporally coherent, and faithful to the underlying manipulation semantics. We evaluate RoboExo on a wrist-to-exo generation benchmark under both seen and unseen settings, where it consistently outperforms all baselines, reducing LPIPS by 76% and FVD by 64\%. Beyond generation quality, RoboExo improves VLA policy learning across six tasks, increasing the average success rate by +30\% with generated exocentric views. These results demonstrate RoboExo as an effective and scalable data engine for enriching low-cost UMI demonstrations. The code will be publicly available.