Cross-Layer Evolution Graph Learning for Fine-grained VLM Hallucination Detection
Abstract
Although Vision-Language Models (VLMs) demonstrate impressive capabilities, their deployment in high-stakes domains is severely hindered by hallucinations. Existing detection methods mostly focus on static internal feature probing with coarse-grained annotations, failing to achieve fine-grained token-level localization without expensive dense labels. In this paper, we reconceptualize VLM hallucination detection as a dynamic graph learning task, identifying hallucinations as distinct trajectory drifts across network layers. Consequently, we propose TRACE (Token Representation Analysis with Cross-layer Evolution), a non-intrusive probing framework that models internal states as heterogeneous spatio-temporal graphs. By employing Graph Neural Networks (GNNs) to capture computational dynamics, TRACE effectively isolates hallucinatory signals. Furthermore, a token-first Multiple Instance Learning (MIL) readout elegantly bridges the granularity gap, distilling macroscopic sentence-level supervision into precise zero-shot token-level localization. Extensive evaluations on M-HalDetect, ViGoR, and HalLoc demonstrate that TRACE achieves state-of-the-art detection performance while being competitive in localization accuracy, offering an efficient and interpretable solution for trustworthy VLMs.