Heads That Write, Not Just Point: Image Retrieval Heads in Vision-Language Models
Abstract
Large vision-language models (LVLMs) can process interleaved text and multiple images in a single context, but how they internally identify the image relevant to a question remains poorly understood. We study this grounding step as in-context image retrieval and introduce ICIR-MCQ, a controlled probe for analyzing image retrieval at attention-head granularity. Using each head's attention mass over image-token spans, we first identify sparse pointer heads that reliably point to the target image. These heads transfer beyond the controlled probe: a single pointer head outperforms external vision-language retrievers on a naturalistic multi-image retrieval benchmark, showing that LVLMs contain a strong internal image retrieval signal. However, attention alone only shows where a head reads from, not what information its output passes on to later layers. We therefore introduce a complementary output-side analysis. For each head, we train a lightweight image retriever on its value-weighted output and use it to identify writer heads, whose outputs contain enough information to identify the target image. Surprisingly, the pointer and writer criteria select substantially different head sets, with only 35 to 44 heads overlapping among the top-100 heads under each criterion. Targeted set-partition ablations show that the heads most important for standard multi-image inference are concentrated in this pointer-writer intersection, while heads selected by only one criterion contribute little beyond random-head controls. These results show that attention-based pointing alone is systematically incomplete for identifying image retrieval heads in LVLMs. Causally relevant retrieval heads must also write target-image information to the residual stream, not merely point to it.