What Remains After Retraction? Evidence Withdrawal in Medical VLMs
Abstract
Multimodal models deployed in clinical workflows accumulate context across a conversation, and some of that context is later invalidated. We study whether an explicit withdrawal of invalidated evidence returns a vision-language model to the behaviour it shows without that evidence. Using chest radiography as the test bed, we evaluate four models (Qwen2.5-VL-7B, MedGemma-4B, GPT-4.1, GPT-5.5) on 388 VinDr-CXR images under seven conversation conditions. Each model first reads an image with a referral note that contradicts the reference label, then receives an explicit withdrawal of that note. Relative to an image-only read, failure after withdrawal is higher in every model, by +8.4 to +53.1 percentage points; a neutral second turn without a note produces no such gap. The residual failure takes different forms: the two open models continue to assert the withdrawn claim, GPT-4.1 predominantly answers uncertain, and GPT-5.5 recovers most but not all cases. A control that withdraws a note consistent with the reference label separates persistence from over-correction and reverses the ranking of two models. A factorial follow-up removes the note text and the model’s own prior answer independently. Retaining the prior answer while deleting the note keeps failure far above the image-only baseline in all three models tested, and for GPT-4.1 exceeds the failure observed when both are retained. Deleting the text of invalidated evidence does not withdraw it. All results are agreement with dataset labels under controlled prompts, not clinical validation.