Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
Abstract
Vision-language-action(VLA) models have emerged as a promising interface for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a ``think with 2D first, drive with dedicated 3D priors'' paradigm: it first identifies sparse 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning, and construct the PlanningGrounding dataset to endow VLM models with planning-oriented grounding ability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of explicit geometric chain-of-thought reasoning for VLA-based planning.