Coarse-to-Refine: Trajectory Self-Refinement in Single Autoregressive Pass for Driving VLA
Abstract
Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving trajectory planning. While explicit test-time thinking has driven mainstream progress in Large Language Models (LLMs), how to effectively utilize reasoning in driving VLA remains unsettled. We revisit this gap through Coarse-to-Refine (C2R), a trajectory self-refinement framework for driving VLA that completes refinement in a single autoregressive pass. C2R first predicts a coarse trajectory, performs explicit reasoning over it, and then generates a refined trajectory autoregressively. We show that the refinement formulation itself drives imitation learning improvement while enabling reinforcement learning to utilize reasoning content for further optimization. C2R establishes VLA state-of-the-art performance across standard and challenging benchmarks, achieving 92.1 PDMS on NAVSIM v1, 90.4 EPDMS on NAVSIM v2, and 49.6 EPDMS on NavHard.