CLEAR: Calibrating Medical Vision-Language Reasoning Under Image Corruption
Koushik Howlader ⋅ Zainab Ghafoor ⋅ Ushashi Bhattacharjee ⋅ Sayantan Chakraborty ⋅ Tanusree Bhattacharjee ⋅ Tirtho Roy
Abstract
Vision-language models (VLMs) are now widely used to answer questions about medical images. A known problem is that they often sound confident even when the image does not support their answer. Earlier benchmarks test whether accuracy drops when the image is blank or wrong, but they miss something just as important: does the model's confidence also fall when the visual evidence is gone? Put simply, is a VLM properly unsure when it cannot really see the image? To find out, we test three open VLMs, including a medically specialized one, on two medical VQA datasets, VQA-RAD and PathVQA, under three image conditions: intact, blank, and mismatched. Alongside accuracy, we report calibration (expected calibration error and reliability curves) and how often the answer stays the same when the image changes. The results are worrying. On VQA-RAD, removing or replacing the image lowers accuracy while mean confidence hardly moves and can even rise, so it nearly doubles the calibration error (ECE $0.23 \to 0.43$). About half of all high-confidence answers become wrong, and 48\% of answers do not change at all when the image is swapped. The effect is even stronger on PathVQA, and, strikingly, it is worst for the medically specialized model, which is the most overconfident of the three. Most importantly, the models notice a blank image more easily than a believable but wrong one, which is the more dangerous case in the clinic. We release our evaluation harness so others can run the same calibration checks on medical VLMs.
Chat is not available.
Successful Page Load