RSPO: Reasoning-Supervised Policy Optimization for Long-Tail Autonomous Driving
Abstract
Vision-Language Models (VLMs) have achieved remarkable advances in general autonomous driving. Nevertheless, effectively exploiting the reasoning capability of Chain-of-Thought (CoT) to enhance the robustness and accuracy of decision and planning in long-tail scenarios still remains a challenging open problem. To this end, we propose a \textbf{R}easoning-\textbf{S}upervised \textbf{P}olicy \textbf{O}ptimization (RSPO) algorithm to boost reasoning and decision-making performance under long-tail driving scenarios. Concretely, we first conduct an empirical analysis of the core challenges faced by current VLMs in autonomous driving decision and planning tasks, and formally define reasoning and decision-making tasks tailored to long-tail scenarios. On this basis, we design a curriculum-guided supervised fine-tuning strategy to enable the model to rapidly fit the mean distribution of model outputs. Furthermore, we propose a reasoning-supervised optimization algorithm, which enhances the robustness and accuracy of reasoning and decision in long-tail scenarios by guiding the optimality consistency between reasoning processes and final decisions. Finally, we construct multiple verifiable autonomous driving reasoning and decision datasets based on open-source benchmarks and conduct extensive multi-dimensional comparative experiments. Experimental results show that the proposed algorithm achieves consistent improvements in prediction accuracy across multiple datasets compared with existing methods. In particular, it achieves a 6.7\% relative improvement over the strong baseline on the CODA-LM task and also yields lower trajectory-prediction errors on downstream planning tasks.