Semantic-Level Invariant Representation Learning for Cross-Hospital Clinical EEG Modeling
Abstract
Robust deployment of clinical electroencephalography (EEG) models requires source-trained models to generalize to unseen hospitals without target-domain calibration, hospital identifiers, or clinical reports at inference. This is challenging because acquisition protocols, hardware, montages, and patient populations perturb low-level EEG statistics, while clinical interpretation is organized around higher-level semantic concepts. We argue that this abstraction mismatch limits conventional signal-level invariant learning: aligning feature distributions alone cannot distinguish clinically irrelevant variation from diagnostic content. We propose Text-Guided Invariant Learning (TG-IL), a deployment-oriented framework that uses clinician-authored reports only during training as semantic anchors for EEG representation learning. TG-IL combines stochastic recording-level aggregation, a variational semantic information bottleneck, and prototype-guided EEG--report alignment to preserve report-derived clinical semantics while mitigating site-specific nuisance variation. At inference, all text-side modules are discarded and the model operates solely on EEG. On a three-hospital clinical EEG benchmark, TG-IL improves zero-shot cross-hospital generalization and worst-case robustness while reducing hospital-specific information in learned representations.