Ground Before You Decide: An Agentic, Clinical-Workflow-Inspired Pipeline for Verifiable Vision–Language Reasoning on Chest Radiographs
Abstract
Multi-modal models can achieve strong predictive performance while relying disproportionately on an easier modality such as text, making it unclear whether individual predictions are actually supported by visual evidence. In clinical settings, reliability and interpretability require knowing what a model predicts and whether the prediction is grounded in clinically plausible imaging findings. For image-centered diagnostic tasks, visual predictions should be explicitly checked for grounding before multi-modal fusion instead of inspecting fused representations post hoc. This principle is instantiated as a diagnostic pipeline for chest X-ray interpretation that mirrors clinical double-reading. Independent vision and text heads produce predictions with supporting rationales, and a reasoning agent checks cross-modal agreement and grounding consistency before accepting a prediction or escalating to human review. Diagnostic pathology terms in the reports are replaced with a generic redaction token to reduce reliance on textual shortcuts. On 200 held-out CheXpert Plus cases across 11 findings, the pipeline answers 82\% of cases, abstaining on the rest, and achieves 95.1\% accuracy on decided cases compared with 84.8\% for text-only and 77.4\% for vision-only baselines.