Iris: Empowering Video MLLMs with High-Frequency Pose Priors via Spatiotemporal Binding
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in general video understanding. Human motion understanding, a dominant topic in video analytics, however, still remains as a bottleneck due to the lack of explicit structural cues in 2D visual patches and the significant motion information loss caused by sparse visual temporal sampling. To address this, we propose Iris (Integrating representations of in-the-wild skeletons), a novel pose-augmented video MLLM tailored for human-centric motion understanding. Iris adopts an asymmetric dual-stream architecture, pairing the traditional sparse vision stream with a lightweight, high-frequency human pose stream to provide spatially structured and temporally dense kinematic signals. To enable pose-vision modality correspondence awareness, we design a cross-modal spatiotemporal binding mechanism, featuring a pose-structured 3D RoPE for precise temporal alignment and a person-grounded cross-attention module for explicit spatial visual grounding. Furthermore, to enable robust large-scale training, we construct an automated pose curation pipeline that organizes pose data into a delicate heterogeneous representation. Extensive experiments demonstrate that Iris achieves superior performance on multiple fine-grained motion benchmarks, while also maintaining highly efficient token computation overhead. Code will be available upon publication.