CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations
Abstract
Multimodal Large Language Models (MLLMs) have shown strong performance on medical public benchmarks, but existing evaluations remain limited as clinically grounded assessments because they often rely on isolated inputs lacking study-level structure and simplified, recognition-style tasks that are far from real diagnostic workflows. We introduce \textbf{CardioLens}, a leakage-resistant evaluation testbed using multi-sequence Cardiovascular Magnetic Resonance (CMR), constructed from private hospital archives via a rigorous report-to-QA construction and verification pipeline. CardioLens is built from 473,896 slices and 13,494 verified QA pairs across 4D Cine, LGE, perfusion, and T2-weighted imaging, and evaluates three stages of CMR interpretation: image understanding, report generation, and disease diagnosis. Across 24 state-of-the-art MLLMs, CardioLens reveals a substantial clinical reality gap. Models perform poorly overall, with performance degrading along the real CMR workflow. Confusion analysis further reveals a category-collapse failure mode: models often default to frequent abnormal categories rather than reliably telling distinct clinical findings. To rule out MLLM-compatible input construction as the failure cause, we compare random, clinically-motivated, and data-driven selection protocols under different slice budgets; performance changes only marginally, typically by about 1\%. Explicit reasoning prompts also fail to rescue performance, often driving models toward conservative predictions rather than improving visual evidence use. These results show that current MLLMs remain far from reliable CMR interpretation, where clinical decisions require integrating distributed evidence across sequences, views, and temporal phases. By exposing where and how current MLLMs fall short, CardioLens provides a clinically grounded testbed for developing the next generation of MLLMs toward real-world clinical depolyment.