CRISP: Compositional Reasoning over Images via Stackable Programs for VLMs
Abstract
How well do vision-language models actually reason about what they see? Current diagnostic benchmarks offer limited answers: they draw from fixed question templates, evaluate only the final answer, and are tied to single visual environments. We introduce CRISP (Compositional Reasoning over Images via Stackable Programs), a platform that generates compositional visual reasoning tasks from reusable, composable operations. Given any scene with a structured specification, the platform constructs reasoning chains of arbitrary depth where each intermediate step has automatically derived ground truth, including bounding boxes for the objects the model should attend to at every stage. New visual environments and reasoning operations plug in without modifying the generation engine. We instantiate the platform across three environments spanning 2D sprites, photorealistic 3D characters, and indoor room scenes, and show that frontier VLMs degrade systematically with reasoning depth and frequently arrive at correct answers through incorrectly grounded intermediate reasoning. Beyond benchmarking, the generated reasoning traces and step-wise ground truth can support future work on training and improving visually grounded reasoning models.