Wait, Predict, or Defer? Evaluating an Offline Sequential Routing Policy for EEG Localization
Abstract
Clinical systems need policies that determine when to act, collect more evidence, or defer to review. We study this decision through a retrospective sequential replay of 4,072 out-of-fold electroencephalography (EEG) segments from 157 patients. The agent receives one 10-second segment at a time and either issues a hemisphere localization, waits for another segment, or defers when it reaches its observation budget. Its primary routing signal, \emph{self-agreement}, measures the fraction of observed segments supporting the modal hemisphere prediction. We compare self-agreement with two other summaries of hemisphere votes, three model-output-magnitude signals, fixed evidence accumulation, and random routing. We select stopping thresholds using training patients and evaluate them on held-out patients across 20 repeated five-fold splits. At a target coverage of 40\%, self-agreement localizes (39.6\%\pm1.1\%) of patients with (0.886\pm0.013) accuracy after (7.51\pm0.66) segments. Vote margin and vote concentration achieve accuracies of (0.879\pm0.015) and (0.874\pm0.016), respectively, while the strongest output-magnitude policy achieves (0.714\pm0.025). Across target coverages of 20\% to 80\%, repeated-vote policies reduce selective risk by 0.086--0.248 relative to maximum output score. The advantage persists across observation budgets and randomized segment orders. These results establish repeated predictions as a stronger routing signal than model output magnitude in this replay. Prospective evaluation should measure how clinician review of deferred cases affects overall system performance.