Post-Deployment Reliability Monitoring of Clinically Deployed LLM-Generated Radiology Indications
Abstract
Large language models (LLMs) are increasingly deployed in clinical workflows, yet post-deployment evaluation remains limited. We study a deployed LLM system that synthesizes longitudinal clinical notes into radiology-relevant indications and introduce MultiJudge, a multi-model LLM-as-a-judge framework for detecting hallucinations, contradictions, formatting errors, and inconsistencies in source clinical data. We validate individual judges and ensemble decision rules against clinician annotation and find substantial heterogeneity across failure categories, with permissive ensemble aggregation increasing sensitivity relative to individual judges. We further evaluate within-model semantic entropy and cross-model semantic disagreement as model-agnostic uncertainty signals for identifying problematic generations. Uncertainty was associated with selected failure modes but varied substantially across generators and was not monotonic with error burden; consistently erroneous generations could remain semantically stable and therefore exhibit low entropy. These findings suggest that post-deployment monitoring should combine explicit quality assessment with uncertainty-based signals, which capture different aspects of clinical LLM failure.