Answerability Is a Missing Safety Axis in Medical Vision-Language Evaluation
Abstract
Medical vision-language models answer fluently when the image in front of them does not match the question. Every widely used medical visual question answering benchmark evaluates only image-question pairs that belong together, so accuracy on these benchmarks says nothing about whether a model knows when not to answer. We argue that answerability is a missing safety axis: matched-pair accuracy is necessary but not sufficient evidence of clinical trustworthiness. We ground the position in an audit of 21 medical benchmarks, of which 11 contain no unanswerable or mismatched condition at all, and in a concurrent controlled study in which generator confidence detects wrong-image pairs at chance (AUROC 0.502 and 0.472), one frozen alignment feature lowers unsafe answering from 0.542 to 0.376, and detection collapses toward chance once the substituted image resembles the original. We propose four requirements for any evaluation offered as evidence of clinical trust: a mismatched condition built from realistic negatives, selective-risk reporting with separate denominators, disclosure of the hardest difficulty stratum, and clinician-anchored labels, and we close with a reporting checklist and a proposal for an open, machine-labeled mismatch benchmark.