Where Vision Controls the Answer: A Causal Account of Grounded Spatial Reasoning in VLMs
Ming (Michelle) Liu ⋅ Liyang Chen
Abstract
Vision-language models (VLMs) often answer spatial questions (is A left of / above B?) correctly, yet where the visual evidence gains causal control over the decision is unclear. Prior work has separately established encoder-side positional sensitivity and decoder-side information pathways, but not the causal link between them. We give a source-valid, orthogonally-controlled causal account of axis-relative left-right and up-down decisions. On geometry-true procedural scenes, a within-scene relation-axis $\times$ intervention-axis factorial localizes the queried relation's control to an early decoder stage (signed difference-in-differences $44.47$ in logit-margin units); this stage is sufficient (a relation-reversing source patch flips the answer; a norm-matched random patch does not) and stage-level necessary (early image-token knockout breaks the answer, the late band mostly does not). Ablating the encoder's width vs. height positional channel gives an axis double dissociation. Restoring the early carrier recovers the answer only when the restored content carries the correct queried-axis relation, a content- and axis-specific encoder$\to$decoder mediation, with an unconditional mediated proportion of $0.90$-$0.99$ across two independent encoders (CLIP, SigLIP) and large language model (LLM) families, while an axis-flipping source is negative (it overshoots to the wrong answer). The early-carrier profile and necessity recur across four VLMs, the mediation reproduces on natural images (What'sUp, ARO) via mirror-flip counterfactuals on two models, and the same early carrier controls two further executable factors (relative size, non-geometric count) on Qwen. Two baselines bound what the protocol establishes rather than compete with it: layer search already matches its site selection, and its advantage over learned steering is architecture-dependent. Rather than a site-selection or steering method, the protocol is a causal diagnostic of whether a VLM's spatial answers are grounded in the image rather than merely correlated with it, applied here to the trials where the model answers correctly.
Chat is not available.
Successful Page Load