Distinguishing Extraction Failures and Clinical Discordance with LLM Uncertainty on Radiology Reports
Jerjes Aguirre-Chavez ⋅ Shant Malkasian ⋅ Roshun Sankaran ⋅ Brendan T Crabb ⋅ Albert Hsiao
Abstract
Large language models (LLMs) are increasingly used to structure information from unstructured text and generate large-scale scientific datasets. Although LLMs can achieve high accuracy in data extraction, extraction failures remain inevitable. As LLM-derived data are increasingly used in downstream research, a critical challenge remains in distinguishing discordance that reflects true differences between datasets from discordance introduced by LLM extraction error. From a real-world cohort of 1,045 patients undergoing coronary CT angiography (CCTA), we evaluated whether LLM uncertainty during vessel-level CAD stenosis extraction could identify extraction failure and clinical discordance. We specifically evaluated patients with expert annotations of CCTA reports (n = 253) or subsequent invasive coronary angiography (ICA; clinical reference standard) (n = 729). LLM uncertainty was measured during the extraction of CAD stenosis severity from CCTA reports. We quantify extraction failure as disagreement between stenosis severity in LLM extracted versus expert annotations, and clinical discordance as disagreement in stenosis severity between CCTA and subsequent ICA. Stenosis Margin, selected from ten candidate measures, was used to quantify LLM uncertainty, and generalized estimating equation logistic regression and AUC evaluated its association and incremental discrimination for each failure mode. LLM uncertainty was strongly associated with extraction failure (adjusted OR 4.15 [95\% CI 2.57--6.69], $P<0.001$) and substantially improved discrimination beyond LLM-extracted CCTA stenosis severity alone (AUC 0.81 to 0.86; $\Delta$AUC $+0.06$ [$+0.03$, $+0.09$], $P=0.001$). In contrast, LLM uncertainty weakly associated with clinical discordance (adjusted OR 1.22 [1.04--1.43], Holm-adjusted $P=0.118$) but provided no incremental discrimination beyond LLM-extracted CCTA stenosis severity ($\Delta$AUC $+0.01$ [$-0.00$, $+0.02$], $P=0.245$). Moreover, extraction errors and clinical discordance showed little overlap, with 38 of 42 clinically discordant vessels (90\%) having been extracted correctly. These findings suggest that LLM uncertainty can identify failures in converting unstructured reports into structured measurements, but that this signal does not necessarily identify discordance in the underlying clinical data. More broadly, failure attribution may be important for evaluating the readiness of LLM-derived scientific datasets for downstream analysis.
Chat is not available.
Successful Page Load