HALMES: Knowing When to Intervene in LVLM Hallucination Mitigation
Abstract
Large Vision-Language Models (LVLMs) have demonstrated strong capabilities in multimodal tasks. Nevertheless, their outputs remain prone to object hallucinations that are inconsistent with visual content. Contrastive decoding, as a training-free inference-time mitigation approach, typically intervenes in the output distribution at every decoding step. However, such indiscriminate full-stage intervention may disrupt originally correct generations and even induce new hallucinations. To address this issue, we first analyze the relationship between internal attention features and hallucinated token segments in LVLMs, revealing that hallucination-related abnormalities already emerge progressively in preceding token segments. Based on this observation, we propose HALMES, a lightweight plug-and-play module that identifies hallucination precursor segments from attention features during generation and selectively triggers contrastive decoding. HALMES can be integrated with various full-stage contrastive decoding methods, transforming indiscriminate intervention into on-demand targeted correction. Extensive experiments across different LVLM architectures and multiple benchmarks show that HALMES further improves the hallucination mitigation performance of existing contrastive decoding methods while reducing the negative impact of full-stage intervention on normal outputs, validating the importance of modeling and selectively intervening on hallucination-prone segments. All data and code will be made publicly available.