Ego-HMB: Human Motion Bridging from Egocentric Images via Motion Bridge Diffusion Model
Abstract
In this paper, we move beyond the conventional setting of motion in-betweening that relies on structured pose inputs, and instead propose a new paradigm that is conditioned on egocentric (first-person) visual observations, thereby reducing the requirements on both users and upstream systems; we term this task Egocentric Human Motion Bridging. This is motivated by realistic consumer scenarios in VR/AR and robotics, where users typically have access only to raw egocentric imagery (e.g., from wearable devices) without reliable human pose annotations or motion capture signals. Unlike conventional motion interpolation in human pose-conditioned settings or motion generation conditioned on a complete temporal language or image context, the egocentric perspective introduces large viewpoint variations, partial body visibility, and severe temporal ill-posedness due to cross-modal endpoint constraints. To this end, we propose Ego-HMB, a unified diffusion-based model that internalizes both endpoint human pose estimation capability and temporally consistent 3D motion infilling ability within a single model through two training stages. Specifically, we first introduce Egocentric-to-Canonical Pretraining (pretraining stage), where we design a hierarchical curriculum with progressively increasing temporal spans. Then, we present a novel Motion Bridge Diffusion Training with endpoint and physics awareness as a fine-tuning stage, where our bridge diffusion diffuse from neighboring poses to enhance temporal continuity while preserving the global constraints between the two endpoints of a 3D motion sequence. Extensive experiments and comparisons with one-stage baselines (a generative model conditioned directly on two egocentric images) and two-stage baselines (estimation models followed by interpolation models) demonstrate that our approach consistently outperforms these baselines.