Verifiable Waveform Reasoning Agents for Clinical Alarm Adjudication
Zhaoliang Chen ⋅ Xiao Hu
Abstract
Clinical alarm adjudication is a high-cost decision problem that requires selecting measurements across noisy waveform channels and applying an asymmetric suppression policy. With tool-using agents increasingly prevalent, it is tempting to delegate both evidence gathering and adjudication to frontier language models. We show that this fails dramatically. On the PhysioNet/CinC 2015 public set, Qwen3.8-Max and DeepSeek-V4-Pro, given the alarm definitions and a deterministic signal-processing toolkit, score 26.52 and 32.86; explicit five-to-one cost guidance raises the best score only to 38.43, below the 39.20 score obtained by never suppressing an alarm. Replacing generated verdicts with one fitted rule rescues frontier performance to 54.39--58.35, revealing a major decision-execution failure, but their acquired evidence still falls short. In light of these failures, we train a compact Qwen3-8B agent in a verifiable environment to elicit disciplined evidence acquisition. Every tool call is stored in a versioned ledger, and the same deterministic rule computes the action from the complete log, so each suppression is replayable without trusting or rerunning the language model. Under record-disjoint, class-stratified five-fold cross-validation over all 750 records, agent-acquired evidence scores 67.28 (95\% CI 62.47--72.16), 8.93 points above the best complete frontier ledger under the identical rule (paired 95\% CI 4.35--13.59), and within 0.34 points of scripted evidence (paired 95\% CI 0.00--0.79). All 356 suppressed alarms have valid run citations and numerically matching evidence. A policy trained only on practice alarms from another hospital also exhibits exploratory cross-site reuse, but its 2.48-point gain over never suppressing is uncertain (paired 95\% CI $-3.85$--9.11).
Chat is not available.
Successful Page Load