In-Context Learning Can Help Vision Language Models Overcome Training Prior
Kun Wang ⋅ Xindi Wu ⋅ Sanghyuk Chun ⋅ Olga Russakovsky ⋅ Esin Tureci
Abstract
Vision-language models (VLMs) learn strong statistical regularities during training, which can make them fail to perceive visual evidence when input images violate those regularities. Such failures are often treated as missing visual capability, but they may instead arise because the model has access to the relevant visual evidence and yet selects an answer dominated by learned training priors. In this work, we use controlled visual in-context learning to show that VLMs can overcome learned priors for visual tasks and use relevant visual evidence and capabilities. Specifically, we show that across five vision-centric tasks, demonstrations yield large gains on prior-conflicting counterfactual examples ($+27.2$% for Qwen3-VL 235B and $+18.1$% for Qwen3-VL 32B) while leaving real accuracy nearly unchanged. These findings also generalize across five VLM families with performance increases from $+9.7$% to $+30.3$%. To explain this effect, we analyze visual representations and find that demonstrations redirect attention toward grounding cues while reducing reliance on cues that support the learned prior. Finally, we find that in-context learning is limited when the underlying visual capability is weak or absent. Together, these results show that current VLMs possess visual grounding capabilities that standard evaluations may overlook, underscoring the need for controlled evaluation design in diagnosing and improving model capability.
Chat is not available.
Successful Page Load