Conviction Without Evidence: Verdict Stability Is Independent of Visual Grounding in Medical Vision--Language Models
Federico Felizzi ⋅ Francesco A Causio ⋅ Bianca D Castaniti ⋅ Olivia Riccomi ⋅ Michele Ferramola ⋅ Vittorio De Vita ⋅ Alessandro Tosi ⋅ Antonio Cristiano ⋅ Alessia Longo ⋅ Lorenzo De Mori ⋅ Chiara Battipaglia ⋅ Melissa sawaya ⋅ Luigi De Angelis ⋅ Marcello Di Pumpo ⋅ Alessandra Piscitelli ⋅ Pietro E Risuleo ⋅ Giulia Vojvodic ⋅ Mariapia Vassalli ⋅ Nicolo Scarsi ⋅ Manuel Del Medico
Abstract
Vision--language models applied to medical question answering can reach correct diagnoses without using the diagnostic image, inferring instead from the clinical vignette. We ask whether this is visible in how \emph{firmly} a model holds its answer: does a model defend a diagnosis derived from a blank placeholder as tenaciously as one derived from a real medical image? Crossing an image-substitution protocol with two levels of adversarial pressure on 60 Italian State Examination items, we find no evidence that ungrounded verdicts are more fragile. In a paired item-level analysis, fabricated consensus (L4) overturns 94.4\% of verdicts with the real image and 98.1\% without ($\Delta = +0.037$; one-sided 95\% upper bound $+0.105$); contentless doubt (L1) overturns 62.5\% and 51.8\% ($\Delta = -0.107$, 95\% CI $[-0.209, -0.001]$), the point estimate running opposite to the direction a grounding detector requires. The two levels nonetheless produce opposite \emph{dynamics}: under L4 the model flips once, always to the exact option argued for (110/110 first-turn flips), and then freezes, averaging 1.9 verdict changes over ten turns; under L1 it never settles, averaging 7.4 changes. Instability under contentless doubt tracks intrinsic item ambiguity rather than evidence availability, and so does not serve as a grounding signal. Flips are almost entirely corrupting (99/100 across both levels). The model is not unaware of the missing evidence---it flags the blank image at baseline on 35 of 60 items and abstains on 4--6---yet those abstentions do not survive contact with pressure. None of the epistemic signals a user could plausibly consult distinguishes a grounded diagnosis from a confabulated one.
Chat is not available.
Successful Page Load