EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries
Abstract
Discharge summaries are crucial clinical documents for patient care, containing clinical context of a patient's overall admission, and are routinely reviewed by medical experts during patient readmission, ongoing care, and diagnostic decision-making. In practice, medical experts often must synthesize information across multiple lengthy discharge summaries by iteratively obtaining information from the documents, while also verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical document question answering, existing benchmarks do not sufficiently capture this clinical setting: they either evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. We construct the benchmark through an expert-informed pipeline that combines a comprehensive discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation. Every single QA sample is then reviewed and revised by 11 medical experts over three weeks. Benchmarking 22 LLMs (both open and closed-sourced) reveals several findings; evidence grounding is consistently harder than content answering, multi-turn errors compound across turns, and strong performance on existing single-turn clinical QA benchmarks does not reliably transfer to this setting. These results establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems.