Looking Is Not Grounding: Better Localized Attention Does Not Ensure Grounded VLM Behavior
Alexander Brady ⋅ Tim Eilers ⋅ Junling Wang ⋅ Mubashara Akhtar
Abstract
Attention maps are widely used as indicators of visual grounding, although it remains unclear whether better localization corresponds to better use of visual evidence in vision-language models (VLMs). These models often ignore the provided image evidence, a behavior reflected in attention maps that frequently focus on irrelevant image regions instead of those supporting the generated tokens. This motivates a central question: does improving attention localization itself improve visual grounding? We investigate this question with a controlled post-training intervention on LLaVA-1.5-7B. We add an auxiliary objective that supervises text-to-image attention using dense grounding annotations, while holding the remaining training setup fixed. The objective increases chance-normalized attention mass on referenced objects by up to $6.8{\times}$ compared to the base model and improves complementary localization metrics, whereas a matched fine-tuning-only control remains near the base model. However, stronger localization does not consistently improve performance across a wide set of grounding benchmarks. Combined with mechanistic analyses of where the supervision is absorbed within the model, our results show that attention localization can be substantially improved without improving grounded behavior, and that better-localized attention maps are not sufficient evidence of greater grounding competence.
Chat is not available.
Successful Page Load