FlatClip: Reusing Image Foundation Models for fMRI Representation Learning via Cortical Flatmaps
Abstract
Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized architectures and costly fMRI-specific pretraining. We ask whether part of this performance gap reflects the importance of preserving cortical geometry. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a training-free surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Because FlatClip is not pretrained on fMRI, it provides an out-of-domain reference for evaluating how much information can be recovered from cortical geometry without relying on fMRI-specific pretraining data. Across resting-state and visual-fMRI benchmarks, FlatClip outperforms ROI/FC-style controls and several general-purpose fMRI foundation baselines, while voxel-level models remain a strong upper comparison. Geometry-control experiments show that disrupting cortical topology reduces performance, whereas mapping ROI-level signals back into atlas-defined 4D volumes partially recovers performance but remains below full voxel inputs. Together, these results position cortical geometry as a key factor in fMRI representation learning and establish surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models.