FlashPlanner: Real-Time Goal-Conditioned Flow-Matching Planning for Autonomous Driving with Online RL Fine-Tuning
Abstract
Diffusion and flow matching have emerged as expressive generative planners for autonomous driving planning, owing to their ability to model high-fidelity and multi-modal trajectory distributions. Nevertheless, existing generative planners are predominantly optimized through imitation learning, which induces a fundamental mismatch between supervised trajectory fitting and closed-loop planning metrics, while their iterative sampling procedures often impose substantial computational overhead. In this paper, we propose FlashPlanner, a goal-conditioned flow-matching planner with online RL finetuning for closed-loop AD planning. FlashPlanner introduces a continuous future goal point as a compact navigation interface for the generative planning policy. This goal is produced by a continuous goal predictor, which is first pretrained and subsequently optimized with RL using closed-loop feedback from multiple candidate goals at each decision step. To support efficient online reinforcement fine-tuning and real-time deployment, FlashPlanner adopts a data-prediction flow-matching objective and removes redundant architectural components in existing diffusion-based planners. Experiments on the closed-loop nuPlan and interPlan benchmarks demonstrate that \textit{FlashPlanner} achieves state-of-the-art planning performance while delivering 6× faster inference (166 FPS) than the previous SOTA baseline (28 FPS). We will open-source our project.