Read but Unused: Dissociating Perception from Readout in Vision-Language Models
Abstract
Vision-language models (VLMs) are often asked to read text that arrives as an image, such as a screenshot, a scan, or a rendered document, and they answer less accurately when they do. When a model gets such a question wrong, accuracy alone cannot tell us why: did it fail to see the text, or did it see the text and fail to use it? We separate these two causes. For Qwen2.5-VL-7B, the problem is not seeing. Asked to transcribe the very images it answered incorrectly, the model reproduces the text with 0.98 median character accuracy. Yet its Yes/No answers drift toward "No," and 89% of its errors on images are cases where it said No when the answer was Yes. Tracing the model's internal activations shows that the difference between a clean render and a corrupted one reaches the position where the model produces its answer in layers 18–20. A single fixed correction to the model's output scores, calibrated on separate control items, recovers 21 accuracy points on a test set built to contain many of these failures (baseline 0.54 by construction). Accuracy on Yes questions rises from 0.42 to 0.72, at a cost of 4 points on No questions. LLaVA-1.5-7B also loses accuracy on rendered text, but for the opposite reason. Its image encoder works at one fixed resolution, so it cannot make out the text at any size we tested (0.04 median transcription accuracy), and its errors look like those of a model that always answers "Yes." The same symptom thus has opposite causes in the two models. A simple transcription check tells them apart, and for the model that can read but not use the text, analyzing its internal activations localizes the problem and partially fixes it without retraining.