Evaluating Transformer Language Models for Emergency Triage Classification in Spanish Clinical Notes
Abstract
Hospital triage based on the Emergency Severity Index (ESI) depends on nursing professionals’ interpretation and application of the triage protocol. In this aspect, introducing subjective and cognitive biases result in a reported mistriage rate of approximately 32% [1]. This study evaluates and compares the performance of four language models based on encoder architecture (mBERT, BETO, RoBERTa-Clinical, and BSC-Bio-EHR-ES), which were fine-tuned on a corpus of 112,052 anonymized real-world clinical notes in Spanish corresponding to the reason for visit of patients treated in the emergency department of the Regional Hospital. The corpus was preprocessed and structured into the five severity categories (C1–C5, ranging from immediate to non-urgent care) with special emphasis on categories C2 and C3, which together account for up to 88% of the total cases and represent a substantial clinical challenge in the event of classification errors (mistriage). The models were evaluated on an independent test set using the following metrics: recall and F1-score (Tables 1 and 2). The results demonstrate that domain-specific models achieved the best overall performance. In detail, RoBERTa-Clinical reached a recall of 0.84 and 0.78 for C2 and C3 category, respectively (Table 1), while both BSCBio-EHR-ES and RoBERTa-Clinical achieved F1-scores of 0.84 for the dominant C2 category and maintained stable performance of 0.79 and 0.80 in C3, respectively (Table 2). Furthermore, classification errors were almost exclusively concentrated in adjacent categories, avoiding multilevel misdiagnoses. In this regard, these results indicate that domain-specific models (RoBERTa-Clinical and BSC-Bio-EHR-ES) can capture clinically relevant distinctions in triage narratives and may provide a useful consistency mechanism for identifying potentially discrepant categorizations. However, the observed performance ceiling likely reflects limitations in the training data, as clinical records inherently contain subjective assessments and potential misclassifications. Consequently, supervised models may reproduce aspects of human variability, rather than reflect an objective clinical ground truth. Nevertheless, specialized language models show potential as clinical decision-support tools due they provide an independent layer of consistency without replacing professional judgment. In this way, future work will focus on a multimodal architecture that integrates structured clinical data (i.e. vital signs and preexisting comorbidities) with unstructured narratives. In this way, this multimodal architecture may enrich the predictive context, mitigating historical biases and improving automated triage reliability. [1]. D. R. Sax, E. M. Warton, D. G. Mark, D. R. Vinson, M. V. Kene, D. W. Ballard, et al., “Evaluation of the emergency severity index in US emergency departments for the rate of mistriage,” JAMA Network Open, vol. 6, no. 3, p. e233404, 2023.doi: 10.1001/jamanetworkopen.2023.3404