Found but Not Bound: Grounded Objects and Unfaithful Relations in Vision-Language Models
Abstract
A vision-language model that answers "is the fan on the ceiling?'' correctly will often accept "is the ceiling on the fan?'' too. The usual explanation is that these models keep the objects a question names and discard its syntax, but that was inferred from answers alone, which cannot say what the model looked at. We ask directly, blurring annotated image regions and recording both the answer and the hidden state behind it. This yields two measures: an exchange index, which places the evidence for a question between two anchors taken from the same image, and an argument-support test, which asks whether a single answer was caused by the objects its question named. Across eight instruction-tuned models the two operations that the bag-of-words account treats as one come apart. Locating the arguments succeeds, under all six removal operators we tried. Binding them to roles is weak: exchanging the arguments flips the correct answer while the evidence moves far less than the answer does. In five of the eight models the largest group of cases is one where the answer still depended on the two named objects and was the wrong way round.