What Entropy Cannot See: Signed Visual Evidence for Training-Free Uncertainty in Vision–Language Models
Abstract
Vision-language models (VLMs) can be confidently wrong when their language prior overrides the visual evidence. This failure is particularly challenging for training-free uncertainty estimation, because many existing scores are functions of the predictive entropies of the with-image and image-ablated next-token distributions, and an entropy does not record the direction in which the image moves the prediction. We show that, under greedy decoding, no score of this form can identify whether the image increases or decreases the probability of the token the model actually emits, and we give in closed form the largest such decrease that leaves every entropy unchanged. The direction is, however, preserved at the emitted token itself. Motivated by this observation, we introduce a score based on the signed realized gain: the change in log-probability contributed by the image to each emitted token relative to a fixed reference. We aggregate these gains into a single scale-free statistic and combine it with two entropy-based views by rank fusion. Our method requires only one greedy decode and two teacher-forced prefills, with no sampling, no auxiliary model, and no trained probe. Across 8 benchmarks and four small-scale models, it outperforms most training-free baselines and matches the remaining ones. At larger scales, up to 38B parameters, it achieves the best performance among all evaluated baselines.