Interpretability nominates a testable hypothesis from an fMRI foundation model: a proof of concept with sparse-feature attribution
Abstract
The search for neuroimaging markers of clinical and cognitive phenotypes has produced many candidates and validated very few, and the difficulty is less a shortage of candidates than the absence of a principled way to decide which of the many quantities derivable from a scan is worth validating. Foundation models trained on large fMRI corpora make that choice implicitly whenever they predict: the prediction rests on particular properties of the signal. Here we ask whether interpretability can recover them, studying BrainLM on the Human Connectome Project Young Adult cohort. We measure what the model prediction relies on, with a supervised distillation into pre-specified signal features and an attribution through an unsupervised sparse-autoencoder dictionary. Both converge on lag-1 autocorrelation in the central executive network: seven such features reproduce 91% of the model's prediction of subject-level fluid intelligence scores. On 280 subjects held out of every step that produced them, the three pre-specified features they nominate match the 512-dimensional model (r = 0.218 [0.103, 0.327] against r = 0.222). Interpretability applied to a brain foundation model yields a candidate for neuroscientists to carry into validation studies.