Focus Matters: Attention-Value Dynamics for Hallucination Mitigation in Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) have achieved impressive progress in multimodal reasoning, yet they remain prone to object hallucinations, generating descriptions of objects that are not present in the input image. In this work, we investigate hallucination from the perspective of attention-value dynamics inside LVLM vision encoders. We identify a consistent three-phase structure of visual processing---diffusion, focus, and rediffusion---and show that the focus phase is where attention most clearly separates strongly and weakly supported visual tokens. However, low attention does not necessarily imply negligible downstream influence: low-attention tokens in the focus phase can still exert non-negligible value-side influence on the attention output relative to their small attention mass. Through controlled phase-wise value interventions, we find that hallucination behavior is particularly sensitive to the value content of low-attention tokens during this phase. Replacing or neutralizing these values reduces hallucination metrics while largely preserving grounded object evidence. A token-level teacher-forcing analysis further shows that the intervention reduces the probability of hallucinated object tokens with much smaller effects on ground-truth object tokens. In addition, Visual Attention Ratio (VAR) analysis shows that focus-phase intervention is accompanied by increased attention to visual tokens during decoding. Based on these observations, we instantiate a simple training-free inference-time intervention that replaces focus-phase low-attention values with an image-level mean value vector using statistics from a single forward pass. Experiments across multiple LVLM backbones demonstrate that this analysis-derived intervention reduces object hallucination with negligible additional runtime and remains compatible with existing inference-time mitigation methods. Our project page is available at: https://anonymous.4open.science/w/FocusMatters-7B32/