Attention-based Routing for Interpretable Multimodal Brain Encoding
Abstract
Understanding how the brain integrates information across sensory modalities under naturalistic settings remains a central challenge in cognitive computational neuroscience. Despite recent progress in multimodal brain encoding driven by the advancement of deep learning, many existing approaches combine representations from pretrained modality-specific models using static aggregation or linear readouts, and therefore do not explicitly capture how information is selectively integrated across cortical regions. Here, we introduce Multimodal Transformer Brain Encoder (M-TBEn), a neural encoding framework that models multimodal integration as a routing problem. M-TBEn extends prior vision-based brain encoding approaches by introducing a cross-attention mechanism for multimodality, in which learnable parcel-level queries dynamically select and aggregate modality-specific representations from visual, auditory, and linguistic inputs. This design is not only parameter-efficient but also yields parcel-specific multimodal representations that support a flexible linear readout for predicting brain responses at various spatial resolutions. We evaluate M-TBEn on naturalistic video-viewing data, including both in-distribution and out-of-distribution stimulus conditions. The model is able to accurately predict parcel-level and voxel-level fMRI responses and generalizes across stimulus distributions. Furthermore, the learned attention patterns provide a structured descriptive basis for analyzing modality integration across cortical parcels, enabling a connection between parcel-specific routing profiles and established functional brain networks. Together, these results suggest that attention–based routing offers a principled computational framework for modeling multimodal integration in the human brain.