Crohn-CoVer: Claim-to-Evidence Reasoning for Auditable Image-Grounded Clinical Reporting
Abstract
Large language models can generate clinically plausible reports from multimodal data. However, plausibility does not guarantee that every statement is actually supported by the evidence required to justify it. This limitation is especially critical in longitudinal medical reasoning, where conclusions may depend on current and prior imaging, the clinical context, and clinical guidelines. We introduce Crohn-CoVer, an evidence-grounded reasoning framework for auditable, image-grounded reporting in longitudinal Crohn’s disease follow-up. Crohn-CoVer formulates verification as a reasoning problem that links each claim to the corresponding evidence: each generated claim is associated with an explicit evidence contract and then mapped to relevant visual, longitudinal, clinical, and guideline-based evidence. The claim is evaluated only when this evidence is deemed adequate. The framework then produces one of three verdicts: \textsc{ENTAILED} (supported by the evidence), \textsc{CONTRADICTED} (contradicted by the evidence), or \textsc{INSUFFICIENT EVIDENCE} (insufficient evidence), accompanied by an evidence trace that enables auditing of the reasoning process. On a controlled benchmark comprising 360 reasoning scenarios and 2,160 atomic claims, Crohn-CoVer achieves a Macro-F1 score of 0.72, compared with 0.52 for LLM-as-Judge and 0.68 for an LLM provided with the same evidence contract. It also reduces the unsafe acceptance of unsupported claims from 0.10 to 0.07. Ablation studies demonstrate the complementary roles of guideline retrieval, evidence adequacy assessment, and claim-conditioned routing. Furthermore, experiments comparing oracle and predicted evidence help distinguish perception errors from reasoning errors. Since the benchmark was constructed under controlled evidence conditions and without independent clinical adjudication, these results primarily characterize the behavior of the verification system rather than its clinical performance. They therefore motivate prospective validation on real patient examinations.