When Medical VLMs Stop Understanding: MedTEC-Bench for Probing Semantic Specificity
Abstract
Medical vision--language models (VLMs) are attracting growing interest as a foundation for clinical image interpretation, retrieval, and decision support. However, current evaluations largely emphasize aggregate accuracy, retrieval, or classification performance, offering limited insight into whether these models possess the fine-grained semantic understanding required for safe and reliable use in high-stakes clinical settings. In this work, we introduce \textbf{MedTEC-Bench}, a diagnostic benchmark for probing semantic understanding in medical VLMs across five levels of increasing specificity, from broad modality recognition to fine-grained findings, and negation. MedTEC-Bench spans five medical imaging modalities: chest radiography, brain MRI, retinal fundus photography, dermoscopy, and histopathology. It combines controlled semantic probes with a suite of metrics designed to reveal failures hidden by aggregate benchmarks. Evaluating a diverse set of VLMs, from broad biomedical models to clinical-specialist models, we find that existing systems often fail on basic yet clinically meaningful semantic distinctions, despite strong reported performance on standard tasks. Our results suggest that current medical VLM evaluations may overestimate vision-grounded clinical understanding and highlight the need for targeted semantic benchmarks before such models are deployed in high-stakes medical workflows.