Grounding Multi-Image Context through Visual Composition
Abstract
Vision-language models often need to relate a query to evidence spread across several images and text segments. We study Visual Context Composition (VCC), a training-free interface that binds related images and text by proximity and places them on one visual canvas. VCC produces its largest and most consistent gains on fine-grained reference matching across six models and also improves peak VL-ICL accuracy. Matched controls separate explicit binding from a residual composition advantage. Explicitly stating image-label relations explains most of VCC's improvement over an interleaved baseline. It does not explain the full improvement: with explicit binding and total raster area approximately matched, VCC retains a 6.1-point average advantage. Synthetic comparisons characterize two associated design factors: routing task-relevant text through the visual encoder and organizing relations through visual grouping. Region interventions then reveal an asymmetric signature of cross-region mixing. On image-query tasks, rendered text is necessary before encoding, but its output tokens are largely dispensable; image-origin tokens, by contrast, remain necessary. These results distinguish explicit binding, visual composition, and evidence of cross-region mixing when fixed multimodal evidence is composed.