MammoGPS: A Benchmark for Visual Grounding, Perception, and Spatial Reasoning in Mammography
Abstract
Vision-language models (VLMs) can achieve strong image-level visual question-answer (QA) performance while relying on shortcuts rather than grounded visual understanding. This is especially concerning in medical imaging, where clinically meaningful interpretation often requires spatial reasoning, localization, and anatomical understanding. Yet in mammography, no benchmark directly assesses whether VLMs can localize findings or reason spatially. To address this, we introduce MammoGPS, a benchmark for evaluating visual grounding, perception, and spatial reasoning in mammography. MammoGPS comprises 56,638 QA pairs across 13 clinical tasks for anatomical and finding grounding, finding-type, quadrant prediction, and depth determination. MammoGPS also includes 94,119 QA pairs across 20 controlled diagnostic tasks designed to probe model failures and spatial biases: we design controlled difficulty ladders to isolate the effects of semantic cues, visual cues, clinical terminology, and background context, as well as spatial-bias tasks to test whether model performance depends on finding location. We evaluate 12 VLMs spanning proprietary, general-purpose, and medical-domain models. Our analysis shows that current systems struggle to localize suspicious findings, make limited use of spatial semantic cues, rely on priors and shortcut heuristics for categorical and grounding tasks, struggle with mammography-specific spatial terminology, and exhibit position-dependent biases. Together, our results underscore the need for more rigorous and fine-grained mammography VLM evaluations.