HIMMEL: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
Haopeng Jin ⋅ Tiankun Yang ⋅ Zhenyu Guan ⋅ ShiQuan Dong ⋅ Wenlong Zhao ⋅ Hongzhu Yi ⋅ Chubin Chen ⋅ Tao Yu ⋅ Jinwen Luo ⋅ Yujia Yang
Abstract
Long video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame count, and weak motion perception under sparse keyframe sampling. Existing remedies either prune visual tokens after decoding, which leaves the expensive RGB pipeline untouched, or they discretise codec motion vectors with a generative tokeniser, which couples motion modelling to a heavy pre-training stage. In this paper we present HIMMEL, a hierarchical video-language framework that allocates semantic and motion capacity along separate paths. A small set of sparse anchor I-frames is routed to the expensive host ViT to ground object identity and scene layout, while the far denser inter-frame intervals are encoded by a lightweight compressed-domain tri-stream adapter that distils motion evidence from motion-vector maps, residual maps, and an I-frame context branch into aligned motion tokens. These tokens are injected into the LLM via a differentiable placeholder mechanism, after a dedicated Stage-1 contrastive alignment that places the motion representation in a geometry compatible with the frozen visual backbone. We further show that an InfoNCE alignment objective beats MSE regression by $+1.5$ pp because it preserves directional motion structure rather than collapsing onto the mean of visual deltas. On Video-MME, HIMMEL surpasses the dense 32-frame baseline by $+2.3$ pp ($61.2 \to 63.5\%$) while using $3.6\times$ fewer context tokens and running end-to-end in $1.03$ s per question instead of $2.75$ s. Switching to a stronger Qwen3-VL-8B host pushes HIMMEL to $64.9\%$, narrowing the gap to much larger proprietary systems while keeping inference on a single consumer GPU. Extensive ablations across stream composition, motion-encoder family, fusion mode, alignment objective, anchor count, LoRA rank, and video duration confirm that the full tri-stream is necessary and sufficient for the observed gains, and that the benefit grows with video length. We release training and evaluation pipelines to support reproducibility.
Chat is not available.
Successful Page Load