BrainVista: Modeling Naturalistic Brain Dynamics as Multimodal Next-Token Prediction
Abstract
Naturalistic fMRI offers a time-resolved window into cortical dynamics during continuous multimodal experience, where future brain states are jointly governed by endogenous neural context and exogenous sensory drive. Forecasting such dynamics, however, remains challenging because fast-changing stimulus streams must be reconciled with the sluggish hemodynamic BOLD response, while cortical activity is organized across functionally heterogeneous brain networks. To address these challenges, we introduce BrainVista, a multimodal autoregressive framework for observed-stimulus conditional brain-state forecasting. BrainVista predicts future fMRI states from the history of brain activity and stimulus tokens temporally aligned to each prediction query, enabling autoregressive prediction without access to future ground-truth fMRI states. This design supports causal, temporally consistent modeling of naturalistic brain dynamics under realistic sensory stimulation. Specifically, we design Network-wise Tokenizers for functionally structured brain representations, a Spatial Mixer Head for cross-network interaction refinement, and Stimulus-to-Brain masking to preserve temporally causal stimulus–brain conditioning while preventing future brain-state leakage and stimulus look-ahead. We evaluate BrainVista on Algonauts 2025, CineBrain, and HAD, three large-scale naturalistic fMRI benchmarks, under long-horizon conditional autoregressive forecasting. BrainVista consistently outperforms baselines, improving pattern correlation by 10.8\% and 9.4\% relative to the strongest baseline on Algonauts 2025 and CineBrain, respectively.