Do Vision Language Models Exhibit Human-Like Attention in Chart Question-Answering?
Abstract
Do humans and vision-language models (VLMs) attend to charts similarly? We compare the Transformer attention of VLMs to human eye gaze during chart question-answering. Linear combinations of VLM attention heads predict average human gaze on par with specialized methods, and individual middle-layer heads correlate with average eye movements, especially in high-performing models. Ablation experiments show these human-aligned heads are causally involved in chart understanding, though not specifically on the examples where their attention matches human gaze. Head-human alignment depends on inter-human agreement: when viewers disagree, VLM attention heads do not capture aggregate or individual gaze patterns as well. Alignment primarily varies with idiosyncratic chart features, rather than question type, plot category, or example difficulty. Overall, VLMs have human-aligned attention heads that contribute to visual reasoning, but they do not consistently adopt human-like strategies for attending to charts, with implications for VLM interpretability and for building human-like artificial chart reasoners.