Teaching VLMs What to Say, Not How to Reason: Rethinking Counterfactual Reasoning in Autonomous Driving
Abstract
End-to-end autonomous driving systems are known to suffer from shortcut learning and causal confusion, limiting robust generalization under distribution shift. To address this, recent Vision-Language Model (VLM) approaches introduce intermediate linguistic reasoning steps including chain-of-thought, chain-of-causation, and counterfactual (CF) reasoning, and report consistent improvements in closed-loop driving. Among these, CF reasoning has been emphasized as enabling self-reflective, high-level causal understanding, yet whether these gains reflect genuine causal reasoning or language-induced distributional bias remains unclear. In this work, we examine what CF reasoning actually induces in VLM-based autonomous driving using accident scenarios with explicit causal structure and a discrete meta-action space. We observe that CF reasoning consistently shifts the action distribution toward more conservative behaviors, accompanied by increased perplexity and policy entropy, and that this effect persists under reinforcement fine-tuning, suggesting a structural rather than transient artifact. However, visual attention analysis and causal intervention experiments show that CF reasoning does not meaningfully alter visual grounding, and VLMs respond similarly to causal and non-causal perturbations, unlike humans who selectively react to causal factors. Overall, these results suggest that CF reasoning does not induce genuine causal reasoning grounded in visual structure. Instead, it acts as a linguistic mechanism that reshapes the action distribution, teaching models what to say but not how to reason, highlighting a fundamental gap between language-based reasoning and causal understanding in VLMs.