Auditing Residualization-Based Evaluation: A Confounding Failure Mode in Cross-Modal fMRI Decoding
Jaden Chun ⋅ Andy Zhang
Abstract
We study a common cross-modal decoding protocol -- residualizing a target modality against a correlated one before evaluation -- and show that, in this fMRI setting, its interpretive prerequisite is not satisfied. Specifically, we train an fMRI decoder using only visual targets and evaluate whether its predictions can retrieve concurrent residualized audio representations from naturalistic movies. The results show that the video target alone can retrieve the audio residual targets at $7.44\times$ chance, indicating that the residualized audio target still contains substantial information accessible from video alone. This result also persists across six orders of magnitude of ridge regularization strength, including under cross-validated regularization. At zero literal temporal overlap in a local candidate-pool stress test, true-video retrieval remains at $4.50\times$ chance. Finally, the brain-derived decoder, which is trained only against video targets, achieves $2.77\times$ the chance in comparison to the $2.04\times$ chance achieved by a validation-calibrated video-mixture baseline and the $1.49\times$ chance achieved by a video-only baseline constructed by adding isotropic Gaussian noise. Overall, these results suggest that while there may be some relationship between the two modalities, the observed effect cannot be attributed uniquely to the neural signals because the same evaluation outcome remains accessible from video alone. Before a residualized target is interpreted as evidence of an isolated construct, the residual should be tested directly for remaining nuisance predictability.
Chat is not available.
Successful Page Load