FusionNeXt: Sequence-First 3D Multi-Modal Fusion in the Era of LLMs
Abstract
Modern LLMs/VLMs have largely converged to a unified architecture: representing heterogeneous inputs as a single token sequence and processing it with a deep stack of generic operators such as attention and feed-forward layers. This sequence-first interface simplifies system designs and thus enables extensive algorithmic and hardware optimization, making it a compelling blueprint for modern model development. In contrast, 3D multi-modal fusion models still rely on complex and specialized designs such as dense feature representations, sparse 3D convolutions, deformable attention, view transformations, etc. In this paper, we ask: \textbf{Can 3D multi-modal fusion embrace the same design path as modern LLMs/VLMs and benefit from their mature and rapidly evolving ecosystem?} We answer this affirmatively and propose \textbf{FusionNeXt}, a modern multi-modal fusion paradigm that (i) unifies camera and LiDAR features into a shared token representation, (ii) serializes tokens into locality-preserving 1D sequences, and (iii) performs feature fusion using a deep stack of vanilla LLM/VLM-style blocks---FlashAttention, pre-norm residuals, SwiGLU FFNs, etc. The proposed paradigm enables standard sequence modeling tools to be applied directly to 3D multi-modal fusion, resulting in fast inference, strong performance, and generalization across tasks. FusionNeXt achieves state-of-the-art (SOTA) results on the nuScenes 3D detection and Occ3D occupancy benchmarks while delivering high inference throughput, suggesting a scalable direction for 3D perception. We will open-source our code.