An In-Depth Analysis of Hallucination Detection Methods for Vision-Language Models
Abstract
Currently, evaluation of hallucination detection methods for vision-language models (VLMs) only reports aggregated statistics, overlooking potential confounding factors and providing limited insight into the strengths and weaknesses of each method. We conduct an in-depth analysis of what drives detection performance, focusing on methods that use either internal model signals (white box) or output logits (black box). First, our analyses show that a naive baseline of a word's token position in the caption is a competitive predictor of hallucinations and that many white box methods, despite performing the best, rely on this heuristic. Second, we find that some white box methods are additionally specialized to VLM architecture, such that when used with certain VLMs, they can even detect hallucinations in captions generated by other VLMs. Lastly, we find that the discriminative advantage of white box methods over black box primarily arises from detecting language-based hallucinations, as opposed to vision-based. Taken together, these analyses reveal insights into hallucination detection methods that are not captured by current evaluation protocols. Drawing from our findings, we encourage future work to develop more comprehensive evaluations that can better reflect hallucination detection behavior.