Look or Read? A Matched Intervention Protocol for Measuring Channel Reliance in Detection-Augmented Driving VLMs
Abstract
Detection-augmented driving pipelines give vision-language models the same evidence twice: as pixels and as text describing detector outputs. Standard accuracy-based evaluation cannot reveal which representation drives a prediction, leaving detector dependence hidden. To identify causal reliance, we propose a matched intervention protocol that toggles visual evidence and detector-derived text independently in a 2x2 factorial design. On matched clips, prompts, and ablation cells, we compare three quantities: performance reliance, the causal effect of each input on the answer; information availability, measured by layerwise probe decodability; and attention allocation at the decision token. The text channel requires localizing the accident-causal road user in every frame; because existing benchmarks lack these bounding boxes, we create, audit, and release them. Applying the protocol to 5 VLMs, 4 prompt pipelines, and two accident benchmarks, we find that the three give different accounts of model behavior. Reliance is set by the question, not by the benchmark or the model. Decodability changes by at most +0.079 where video has its largest causal effect (+0.31), while attention reallocation reveals almost nothing about causal reliance. Layerwise analysis suggests a mechanism: textual information becomes decodable within the first tenth of the network, whereas the same information from pixels continues resolving through the final layer. Models then favor whichever representation reaches the answer in fewer steps.