From Clever Hans to Hypotheses: Interpreting EEG Foundational Transformers
Abstract
Emerging foundation models ( FM s) in electroencephalography ( EEG ) promise a path for deep learning in diagnostics and brain-computer interfaces despite data scarcity, yet their opaque nature remains a barrier to wider adoption. We investigate attention-aware Layer-wise relevance propagation (LRP ) as a post-hoc attribution method for EEG- FM s, extending its use on convolutional neural network ( CNN)- based EEG models to EEG- FM s. We find that LRP can both verify EEG - FM decisions and surface novel, biologically plausible hypotheses from them. In motor imagery, it unmasks "Clever Hans" behavior where models prioritize task- correlated ocular signals over the intended motor correlates. In a naturalistic paradigm for affect prediction, it reveals a recurring reliance on a central electrode cluster, suggesting a candidate sensorimotor signature of arousal. Despite heatmap ambiguity, our results position LRP as a tool for verifying and exploring EEG-FMs as these models mature.