Multimodal Context-Aware Human Motion Generation with Language, Vision, and Object
Abstract
Motion generation has made substantial progress in synthesizing human motion from language, yet text alone remains ambiguous for specifying fine-grained spatial, temporal, and interaction details. Visual observations offer a powerful source of complementary context: even sparse images can reveal intermediate body states, object affordances, and interaction relations that are difficult to specify precisely in language. Whatmore, contextual constraints like visual, object, and partial-motion conditions can ground motion genetration with higher physical fidelity and controllability. We therefore study a unified multimodal context-aware framework for motion generation under diverse conditions. The key challenge is that heterogeneous conditions impose constraints at different spatial-temporal granularities and exhibit token- and phase-dependent relevance. Moreover, incorporating multiple modalities without a robust motion prior may entangle modality-specific semantics and cause negative transfer. In these regards, we combine global multimodal modulation with fine-grained motion token-level context selection, enabling the model to adaptively exploit relevant signals during generation. We further adopt progressive training that first learns a strong language-conditioned motion prior, then extends to multimodal context with different modality combinations. To support multimodal training, we introduce Mo900H, a large-scale benchmark integrating 21 motion datasets with over 900 hours of human motion. Our method reduces FID by 21.1% and 48.3% compared with SOTA methods on the HumanML3D and Mo900H datasets, while also improving motion captioning and enabling diverse multimodal-conditioned generation.