Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation
Abstract
Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task due to the extreme viewpoint changes, extensive occlusions and unobserved regions, where standard geometric conditioning becomes highly unreliable. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which renders a 3D reconstruction into the ego view for overall scene layout, with a semantic grounding branch, which anchors per-object context onto the scene geometry so that heavily distorted regions and unobserved regions can be hallucinated in context. Additionally, to support robust training at scale, we substantially improve the labeling pipeline for real-world data and further develop a fully automated synthetic data engine that renders high-quality dynamic humans in procedurally generated environments. Evaluations on the challenging EgoExo4D dataset show that our method substantially outperforms recent state-of-the-art approaches. Extensive ablations further demonstrate that both structured semantic grounding and rigorous data processing are essential for robust exo-to-ego video generation in extreme scenarios.