A Unified Image and Video Encoder for Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) have rapidly advanced machine intelligence by integrating visual perception with powerful linguistic reasoning. However existing visual backbones predominantly encode videos in a rigid per-frame manner and delegate temporal modeling entirely to the language model. This paradigm fundamentally constrains long-range temporal understanding and imposes severe computational inefficiencies. To resolve these bottlenecks we propose FuxViT, a unified Vision Transformer natively supporting a 64K-token context window to seamlessly process diverse inputs ranging from static images to hour-scale videos. By employing a joint spatiotemporal pretraining paradigm, FuxViT generates highly coherent and information-rich visual representations. This strategic early fusion frees the downstream language model from reconstructing low-level temporal structures to focus entirely on high-level reasoning. Furthermore we introduce Frame Selective Attention (FSA) as a native sparse routing mechanism. FSA dynamically attends to the most relevant historical frames to achieve exceptional computational efficiency without sacrificing fine-grained spatial or temporal details. Extensive experiments across 16 image and video benchmarks demonstrate that FuxViT consistently outperforms existing baselines and establishes itself as a vastly superior alternative to standard static ViTs for multimodal understanding. We will fully open-source all model checkpoints.