InfiniteVL: A Systematic Approach to Highly-Efficient, Ultra-Long Multimodal Understanding
Hongyuan Tao ⋅ Bencheng Liao ⋅ Shaoyu Chen ⋅ haoran yin ⋅ Qian Zhang ⋅ Wenyu Liu ⋅ Xinggang Wang
Abstract
Processing ultra-long multimodal inputs efficiently remains a critical bottleneck for Vision-Language Models (VLMs). This systematic challenge is essentially three-fold: (1) \textbf{Efficient Architecture}: compressing redundant visual sequences without sacrificing fine-grained details; (2) \textbf{Knowledge Transfer}: seamlessly migrating the capabilities of strong pretrained VLMs into more efficient novel architectures; and (3) \textbf{Generalization}: maintaining robust perception and reasoning across diverse multimodal domains. To address this, we introduce \textbf{InfiniteVL}, a linear-sparse hybrid VLM framework designed for efficient long-context understanding. At its core, InfiniteVL pairs a linear architecture to compress long-range visual context with sparse attention to retrieve precise local details. To support this architecture in moderate academic resources, we developed a three-stage knowledge transfer pipeline backed by capability-oriented data construction instead of training from scratch. This ensures smooth architectural alignment, rapid capability recovery, and robust adaptation across diverse domains. Furthermore, we adapt this foundation to specific deployment scenarios, deriving Sparse InfiniteVL for long-video analysis and Streaming InfiniteVL for real-time scene perception. Extensive experiments demonstrate that InfiniteVL matches the performance of leading VLMs while achieving a 1.7$\times$ decoding speedup and a 2.7$\times$ reduction in memory. Notably, Sparse InfiniteVL accelerates prefilling by 5$\times$ at a 256K context length, while Streaming InfiniteVL delivers stable 25 FPS real-time processing with a constant memory footprint.
Chat is not available.
Successful Page Load