MED-REFLEX: Contradiction-Driven Self-Correction for Multimodal Medical Agents
Abstract
Multimodal clinical reasoning requires integrating evidence from chest X-rays, radiology reports, laboratory data, and prior studies. These sources frequently contradict one another—a report may describe a clear lung while the imaging shows bilateral opacities, or elevated inflammatory markers may provide evidence that is discordant with a benign radiological interpretation. Existing medical AI agents either ignore such contradictions or apply generic self-critique without localising their source or verifying them with targeted evidence. We introduce MED-REFLEX (Multimodal Evidence Detection and REsolution via FLEXible contradiction-correction), an agent that treats evidence contradiction as an explicit state variable and triggers a structured Contradiction Resolution Loop (CRL): detect, classify, localise, verify, revise, re-check. We also introduce MedCon- flictBench, a curated benchmark of 1,847 cases across seven evidence-conflict categories derived from MIMIC-CXR and MIMIC-IV. We define six new eval- uation metrics—Contradiction Discrimination Accuracy (CDA), Contradiction Localisation Accuracy (CLA), Contradiction Resolution Efficiency (CRE), Self- Correction Gain (SCG), Correction Faithfulness Score (CFS), and Over-Resolution Rate (ORR)— to measure whether agents can not only resolve contradictions but also distinguish genuine conflicts from clinically compatible disagreements. On the linked MIMIC-CXR/IV cohort (48,921 studies overall, eight thoracic pathol- ogy classes (Atelectasis, Cardiomegaly, Consolidation, Edema, Pleural Effusion, Pneumonia, Pneumothorax, and No Finding)), MED-REFLEX achieves 87.2% accuracy, AUROC 0.921, ECE 0.058, outperforming multi-agent debate by 1.8 pp AUROC, while using fewer tool calls (3.4 vs. 7.3 for debate) and achieving CDA of 0.891 against 0.773 for the best baseline.