Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models
Abstract
Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning failure, or use simplistic synthetic scenes lacking the realism modern VLMs are tuned for. We introduce Auto-Comp, a fully automated, concept-driven pipeline that bridges this gap by generating photorealistic compositional benchmarks at scale. Its core innovation is a parallel A/B construction: for each concept, the pipeline emits a Minimal sample (template caption, isolated objects on a white background) and a Contextual sample (LLM-rewritten caption, objects embedded in a realistic scene), isolating core binding ability from visio-linguistic complexity. We instantiate four task families spanning the two canonical axes of compositional binding: Color and Shape-Color (attribute binding), and Position and Relative Size (relational binding). We evaluate over 25 VLMs spanning CLIP, SigLIP, hard-negative-trained, and frontier generative models. The findings are consistent across architectures and scales: every model exhibits a large Swap-vs-Confusion gap, with low-entropy distractors (e.g., repeated objects or colors) exposing failures beyond the known bag-of-words limitations. We further uncover a task-dependent trade-off: visio-linguistic context aids relational reasoning but hinders attribute binding through visual clutter. We publicly release the pipeline and benchmarks.