Calibrating Agentic LLMs for Clinical Prediction
Zheqi Lv ⋅ Fuying Ding ⋅ Meihan Wei ⋅ Haoyang Li ⋅ Christos Faloutsos ⋅ Fei Wang ⋅ Chengxi Zang
Abstract
Agentic Large Language Models (LLMs) show great promise in clinical prediction using real-world clinical data. While recent studies have primarily examined their discrimination performance, their calibration performance—that is, whether predicted probabilities accurately reflect absolute clinical risk—remains largely unknown. This gap limits their use in high-stakes medical settings, where well-calibrated risk estimates are essential for guiding treatment and prevention decision-making. In this work, we provide theoretical analyses of how agentic LLM architectures, task knowledge, and tool using affect calibration performance in clinical prediction. Motivated by these analyses, we propose ${C^2Align}$, a Calibrated and Clinically Aligned framework, to optimize calibration while maintaining or even improving discrimination. ${C^2Align}$ aligns two complementary agent branches: an ensemble of task-homogeneous LLM agents and a set of clinical task-inspired LLM agents. This alignment is achieved through a novel discrimination- and calibration-aware loss that encourages clinically meaningful agreement while controlling probabilistic miscalibration. We evaluate this framework on four clinical prediction benchmarks—spanning prognostic prediction and disease diagnosis tasks—using diverse LLMs, agent architectures, and calibration methods over both MIMIC-III and IV datasets. Extensive experiments, supported by theoretical motivation under assumptions, demonstrate that ${C^2Align}$ improves calibration while maintaining or enhancing discrimination. Our work provides a theoretical, methodological, and benchmark reference for developing well-calibrated agentic LLMs for trustworthy clinical prediction.
Chat is not available.
Successful Page Load