Voxel as Token: A New Perspective for Zero-shot Cross-subject Vision Decoding
Yulong Liu ⋅ Ziqiu Huang ⋅ Hua Xu ⋅ Guibo Zhu ⋅ Sirui Han ⋅ Yike Guo
Abstract
Recent advances in brain decoding have achieved impressive visual reconstruction under subject-specific settings, yet they fail to generalize to unseen subjects without retraining---hindering practical zero-shot cross-subject applications. The core challenges stem from inter-subject variability in voxel dimensionality and misaligned spatial response patterns. While existing methods rely on surface-based alignment or complex feature disentanglement, they often operate as black boxes and overlook the potential of volume-based fMRI data. To address this, we propose \textbf{Voxel as Token (VoxTok)}, a conceptual framework that treats each fMRI voxel as an independent token endowed with a learnable functional embedding. This formulation naturally handles variable-length inputs and subsumes existing adapter-based models as special cases where linear projections approximate voxel functionality. Guided by the hypothesis that large-scale spatial distributions of neural activity are consistent across subjects despite fine-grained variability, we introduce a KNN-based estimator to predict functional embeddings for unseen subjects. This approach achieves state-of-the-art zero-shot retrieval performance and comparable reconstruction quality, while enabling plug-and-play adaptation for existing models (e.g., MindEye2). Integrating multi-scale ROIs further boosts performance, with our best model achieving $40.4\%$ zero-shot image retrieval accuracy, rivaling early supervised methods. Crucially, we find that decoding success correlates strongly with ROI size, indicating that global co-activation patterns drive cross-subject generalization.
Chat is not available.
Successful Page Load