Do Robust LVLMs Hallucinate More? Uncovering Robustness-Induced Hallucination in Large Vision-Language Models
Md Zarif Hossain ⋅ Awal Ahmed Fime ⋅ Ahmed Imteaj
Abstract
Robustness and accurate alignment with visual content are fundamental to the reliable deployment of large vision-language models (LVLMs). In this work, we uncover a previously overlooked trade-off: improving adversarial robustness can inadvertently increase hallucination in LVLMs. Through extensive evaluations on both open-ended generation and discriminative benchmarks, we reveal that adversarially trained LVLMs consistently produce outputs that are less aligned with the visual input, despite their improved robustness. To understand this trade-off, we analyze the visual token embeddings of robust encoders and find that adversarial training increases inter-token similarity by up to $38$\%. We refer to this phenomenon as spatial token homogenization. We further demonstrate that spatial token homogenization degrades visual information and amplifies decoder's reliance on language priors. Motivated by our findings, we propose Homogenization-Aware Latent Steering (HALS), a training-free inference-time method that steers decoder hidden states away from homogenization-induced hallucination without modifying the robust vision encoder. Extensive experiments across five hallucination benchmarks show that HALS consistently outperforms existing hallucination mitigation methods on robust LVLMs, reducing CHAIR$_S$ by $31.3$ points on average, improving POPE accuracy by $2.6$ points and achieving a $30.6\%$ reduction in hallucination on AMBER. Moreover, robustness evaluations confirm that HALS preserves adversarial performance, demonstrating that robust LVLMs can be made visually grounded without sacrificing their performance.
Chat is not available.
Successful Page Load