Rejected When Asked, Followed When Assumed: False Presuppositions in Medical Vision-Language Models
Abstract
Many clinical questions assume a finding is present, for example by asking where it is or how extensive it is. We test whether a model that judges a finding absent when asked directly will still describe it when a separate question assumes it is present. Across five open models, 82.9% to 100% of absent findings judged absent under the direct question are nevertheless treated as present under presupposition. A frontier reasoning system does so in 72.2% of cases. The open models still answer at least 95.5% of matched questions when the finding is present. The behavior also remains frequent even when the model is highly confident that the finding is absent, across alternative question forms, and on clinician-authored high-risk cases. We next test whether the behavior can be reduced. Asking the model to verify first helps only partly, while telling it whether the finding is present has a much larger effect. Hallucination-focused fine-tuning can also reduce the behavior, but it shifts the model's answers to direct presence questions at the same time. A lower false-presupposition error rate therefore does not by itself show that the model has become better at telling whether a finding is present.