When Audio Residuals Are Still Visible: A Failure Mode for Interpretability-Driven Discovery in fMRI
Abstract
Interpretability-driven scientific discovery requires evaluations that distinguish a proposed scientific construct from alternative routes to the same observable. We examine this requirement in a cross-modal fMRI decoding protocol that residualizes one correlated modality against another before evaluation. An fMRI decoder is trained only against visual targets and then tested on its ability to retrieve concurrent residualized audio targets. Under the originally specified ridge residualizer, true video alone retrieves the residualized audio targets at 7.44× chance. The failure persists under validation-selected ridge (3.74×) and a nonlinear residualizer that achieves higher validation R² in all 20 folds (4.58×). In a local candidate-pool stress test reaching zero literal candidate–target overlap, true-video retrieval remains at 4.50× chance. Brain-derived predictions retrieve the same targets at 2.77× chance, compared with 2.04× for a validation-calibrated video-mixture control and 1.49× for a Gaussian control. These findings do not establish that the brain-derived effect is stimulus-driven. Rather, they show that a stimulus-only route to the evaluation outcome remains available, preventing a unique interpretation as audio-related neural information beyond the concurrent visual stimulus. More broadly, transformed targets used for interpretability-driven discovery should be audited directly for residual nuisance accessibility under the same representation and evaluation metric used to support the scientific claim.