Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Abstract
A small, invisible change to an image can fully corrupt what a vision-language model's visual encoder represents internally, while the model's actual output stays correct. We call this the train/inference gap and explain it mechanistically on Qwen2.5-VL-7B-Instruct across 200 COCO images. Pixel-level image properties -- texture, frequency, a known attackability measure from CNN research -- cannot predict which images get corrupted. Looking inside the model, we find the attack makes real progress at the exact point that matters: it pushes the correct next word's rank down by roughly ten times, but never far enough to actually change what the model says, and how much progress it makes depends systematically on the same image groups we identify elsewhere in the paper. Tracking this signal through all 28 layers of the language model shows the visual encoder corrupts every image by a similar amount, but the decoder then treats images very differently -- amplifying the corruption for some, and actively pushing it away for others. The real decision about whether an attack works happens in the language model's decoder, not its vision system, which has direct consequences for where robustness defenses and faithfulness checks for deployed vision-language systems should actually be placed.