ALLMTS: Automated LLM-based Medical Triage System
Abstract
Telephone triage is a safety critical task: from a brief intake note, a nurse must decide whether a caller needs an ambulance, an emergency department, an appointment, or self care at home. Queried directly, large language models (LLMs) are poorly calibrated for it: they over-escalate routine cases and cannot justify their recommendations against an auditable clinical standard. We present ALLMTS (Automated LLM-based Medical Triage System), which grounds every disposition in the same commercially maintained telehealth guidelines nurses use. Rather than asking an LLM to triage, ALLMTS makes it \emph{execute the guideline}: it extracts the patient's demographics and reasons for visit, selects the appropriate guideline, evaluates the guideline's triage assessment questions and derives the disposition deterministically from the first positive question - the same traversal rule a nurse follows. On 566 expert validated scenarios spanning 14 dispositions, ALLMTS achieves 95\% accuracy, compared with 47\% for a raw-LLM baseline. On a curated 50-scenario subset released as an open benchmark, ALLMTS reaches 98\% accuracy, on par with the 94.4\% correct disposition rate that a five nurse expert panel recorded on those same scenarios. Critically, on that subset the baselines over-triage up to 50\% of cases (44\% on the full dataset), whereas ALLMTS reduces over-triage to at most 2\% while maintaining low under-triage. These results suggest that coupling LLM language understanding with executable clinical guidelines, rather than relying on parametric knowledge or generic retrieval, offers a practical route to automated triage.