From Clinical Narratives to Research-Ready Data: External Validation of Local LLM Extraction and Prognostic Utility in Heart Failure
Shinnosuke Sawano ⋅ satoshi kodera ⋅ Takenobu Shimada ⋅ Risa Kishikawa ⋅ Shun Kitamura ⋅ Eisuke Amiya ⋅ Junichi Ishida ⋅ Daiju Fukuda ⋅ NORIHIKO TAKEDA
Abstract
Post-discharge risk models for heart failure (HF) rely largely on structured data, while clinically meaningful information in discharge summaries remains underused. We developed an end-to-end, on-premise large language model (LLM) pipeline to extract eight prespecified clinical findings without transferring privacy-sensitive text outside the institution. We evaluated extraction against physician review internally and at an independent external site, compared identical multivariable survival models using physician-reviewed versus LLM-extracted variables, and assessed the prognostic value of the eight findings with 14 structured covariates for two-year all-cause mortality and HF hospitalization in 1,654 patients. Internal agreement was 95.8%, with finding-specific Cohen’s $\kappa$ values of 0.76–0.96, and external agreement was 93.9% ($\kappa = 0.76$). Effect estimates were directionally consistent in 15 of 16 associations, with a median absolute difference in log hazard ratios of 0.10. Adding the findings improved discrimination for all-cause mortality from 0.696 to 0.726 ($\Delta C = 0.030$, 95% CI, 0.015–0.056) and for HF hospitalization from 0.633 to 0.671 ($\Delta C = 0.038$, 95% CI, 0.025–0.056). The on-premise LLM converted clinical narratives into interpretable structured variables with high agreement with physician review, broadly preserved downstream effect estimates, and improved prognostic performance beyond structured data. These findings support a privacy-preserving, scalable approach to converting clinical narratives into research-ready data.
Chat is not available.
Successful Page Load