Scaling fMRI Foundation Models with Native-Space Data from Heterogeneous Public Studies
Xujin C Liu ⋅ Connor Lane ⋅ Hassan Muhammad ⋅ Tanishq M Abraham ⋅ Paul Scotti
Abstract
Preparing fMRI for foundation model training usually involves computationally expensive data cleaning and alignment steps, including nonlinear deformation of brains to a shared anatomical template. We test a minimalist approach that preserves brains' native shape and greatly simplifies the preprocessing pipeline, placing more burden on the foundation model. Our pipeline prepares the average participant's data in less than $16$ minutes on one CPU core, enabling us to train the first fMRI foundation model on nearly the entire corpus of OpenNeuro: $22{,}373$ recording hours from $837$ studies and approximately $30{,}000$ participants. We evaluate the resulting model on seven classification tasks on the standardized Brainmarks benchmark suite. Our hypothesis was that native-space pretraining would suffer compared to traditional template-space at matched data scale, but that we would close the gap by scaling to more data. Surprisingly, however, we find that native-space pretraining performs roughly on par with template-space, and that scaling to larger data provides only modest further improvement. Our work represents an initial step toward large-scale native-space pretraining, but more work is needed. We release our preprocessing code and training-ready data shards to facilitate future work in this direction.
Chat is not available.
Successful Page Load