EIHMR: Collaborative Human-Camera Estimation for Global Human Mesh Recovery
Abstract
Recovering global 3D human motion from monocular video captured by a moving camera is a fundamental yet challenging problem, as camera ego-motion and human body motion are tightly entangled in the image observations. The prevailing two-stage paradigm treats camera estimation and motion reconstruction as isolated processes, causing errors on both sides to be further amplified when combined in the world coordinate system. To address this, we draw inspiration from the human inner visual simulation mechanism and propose EIHMR, a collaborative human-camera co-estimation framework. EIHMR comprises two complementary modules that bridge scene-aware human motion refinement and motion-aware camera estimation: Scene-Aware Local Human Motion Reconstruction reprojects motion sequences into frozen keyframe viewpoints and leverages metric depth and kinematic constraints to produce geometrically consistent local motion, while Motion-aware SLAM re-renders the refined motion as static meshes in the original frames, converting dynamic human regions into structured matching cues for robust camera estimation. EIHMR consistently improves global trajectory reconstruction over strong baselines, demonstrating the effectiveness of collaborative human-camera estimation for long-range human motion recovery.