Are EEG Foundation Models Worth It? A Controlled, Multi-Seed Audit
Abstract
EEG foundation models report downstream gains under protocols that rarely include a matched random-initialisation encoder, a band-power baseline on the same split, or a layer chosen on validation. We evaluate five public-checkpoint EEG foundation models and an in-house JEPA-style model under one protocol with all three controls and multiple seeds, on five clinical and motor-imagery benchmarks plus event-related-potential, sleep, emotion and within-subject motor-imagery tasks. Choosing the probed layer on the test metric inflates the layer-wise peak by +0.034 on average and reverses the peak-versus-output gap in 9 of 36 cells. Under a validation-selected frozen probe, random initialisation matches or beats pretraining for half of the models on sleep and on emotion, and a five-band logpower baseline beats every model on emotion and on motor imagery. Under full fine-tuning with five seeds, 5 of 30 cells show a positive gain beyond twice the seed std, and the one reliably negative cell is a loss of −0.290 AUROC that survives three matched learning rates. We distil four reporting practices and release every cell for re-analysis.