H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
Abstract
Parametric human models capture global pose but overlook fine-grained surface dynamics. Generic scene flow estimates dense motion but struggles with articulated humans and lacks 4D ground truth. To bridge this gap, we introduce H-Flow, a dense 4D motion representation that captures non-rigid deformations beyond skeletal kinematics. We estimate H-Flow from monocular video via a unified multi-head architecture that jointly perceives pose, depth, and flow. To overcome the scarcity of dense motion annotations, we propose a self-supervised learning paradigm that embeds geometric, structural, and biomechanical priors into optimization. These cross-modal constraints tightly couple pose, depth, and flow, so that improving any one modality simultaneously drives the others toward consistency. For evaluation, we present DynAct-4D, a high-fidelity synthetic benchmark providing dense 4D ground truth for complex human movements. Extensive experiments show that our method outperforms state-of-the-art scene flow baselines. Furthermore, H-Flow serves as an effective motion primitive, bringing substantial improvements to downstream tasks including action recognition and video generation. Models, code, and resources will be released upon publication.