FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery
Abstract
Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively well constrained, whereas distal articulations such as the arms and legs are more uncertain. Building on this observation, we propose \textbf{FactorizedHMR}, a two-stage framework that treats these two regimes differently. A deterministic regression module first recovers a stable torso-root anchor, and a probabilistic flow-matching module then completes the remaining non-torso articulation. To make this completion reliable, we combine a composite target representation with geometry-aware supervision and feature-aware classifier-free guidance, so that well-observed body parts remain accurate while ambiguous articulations can be completed without collapsing to deterministic averages. We also introduce a camera-aware synthetic data pipeline that provides the paired image-camera-motion supervision under diverse viewpoints. Across camera-space and world-space benchmarks, FactorizedHMR improves articulation recovery and reduces drift relative to strong baselines, especially in ambiguous scenes.