Where Are MLLMs Looking When They Hallucinate? Mitigating Visual Hallucination via Gaze Steering
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress, yet visual hallucination remains a critical barrier to reliable deployment. This paper aims to answer a fundamental question: when MLLMs hallucinate, where are they actually looking? By tracing the inter-layer consistency of visual attention, we reveal that hallucinated reasoning trajectories bifurcate into two pathological states: Attention Locking, where the gaze becomes overly focal and rigidly anchored to limited visual evidence, and Attention Collapse, where attention becomes excessively dispersed and fails to ground reasoning in meaningful cues. In contrast, faithful reasoning maintains a dynamic equilibrium between established visual evidence and broader peripheral exploration. Building on this insight, we propose ReGaze, a training-free framework that redirects the model's gaze toward this healthy equilibrium to mitigate hallucinations. Specifically, ReGaze monitors layer-wise gaze consistency and triggers timely interventions whenever an unhealthy gaze state emerges. For the locked state, ReGaze redistributes attention from over-dominant critical tokens to peripheral regions, encouraging the model to explore richer visual details. For the collapsed state, it gathers scattered attention from non-critical tokens and injects it into critical anchors, amplifying key visual signals. Experimental results demonstrate that ReGaze enables MLLMs to process visual information more effectively and significantly mitigates hallucinations.