MedVIGOR: Visual Evidence Internalization for Observation-Driven Reasoning in Medical VLMs
Abstract
Medical image interpretation requires diagnoses grounded in case-specific visual evidence. However, medical vision-language models often produce plausible answers by exploiting clinical language priors and report-level co-occurrence patterns rather than faithfully using the input image, leading to Visual De-anchoring and Spatial-Semantic Binding Breakdown, where reasoning detaches from visual evidence or binds correct answers to incorrect anatomical support. In this work, we introduce MedVIGOR, a framework that internalizes visual evidence as an intrinsic constraint on medical VLM reasoning. MedVIGOR combines visual-necessity supervision, dual-pathway visual anchor internalization with MedSAM-derived spatial support and DINOv2-derived semantic patterns, and evidence-consistent reinforcement learning to preserve alignment among visual evidence, intermediate reasoning, and final answers. We further present MedVIGOR-Bench, a unified benchmark for evaluating whether medical VLMs are both diagnostically correct and evidentially grounded across concept recognition and visual localization tasks. Experiments show that MedVIGOR improves diagnostic accuracy and grounding fidelity over strong medical VLM baselines, highlighting the value of internalized visual evidence for trustworthy medical reasoning.