OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
Abstract
Commercial egocentric devices capture human behavior in everyday settings, yet full-body motion reconstruction must generalize across diverse cameras and mountings. This is challenging because head trajectories under-determine body motion, while hand observations are intermittent and camera-dependent: a hand may disappear simply because it has left the camera's field of view. Existing methods treat such absences as missing data, relying on fixed visibility regimes or post-hoc optimization, and fail to generalize across devices.We present \modelname, a sequence-level diffusion framework for device-agnostic egocentric full-body motion reconstruction. Our key insight is that egocentric videos contain sequence-level evidence about both the wearer and the camera-induced visibility structure, enabling \modelname to infer consistent body shape and coherent motion from sparse head and hand cues. To prevent overfitting to a single camera setup, we further introduce geometry-aware visibility augmentation, which synthesizes visibility patterns from realistic variations in camera geometry. We also present OmniEgoDB, the first mocap benchmark with ground-truth motion captured across multiple consumer egocentric devices. Experiments on synthetic, real-device, and in-the-wild settings demonstrate superior motion reconstruction, robust cross-device generalization, and coherent in-the-wild results.