RAPDrive: Shared-Latent Hybrid Decoding for Reasoning and Planning in Autonomous Driving
Abstract
Driving vision-language-action (VLA) models must connect semantic reasoning with executable trajectory planning, but language and motion exhibit different generation structures: reasoning text is naturally sequential, whereas future trajectories require horizon-level geometric consistency. We present RAPDrive, a shared-latent hybrid-decoding framework that processes visual context, ego state, motion history, reasoning tokens, and future planning slots within a single transformer sequence. RAPDrive generates reasoning text autoregressively while refining future motion through iterative masked denoising in a dedicated motion-token space, followed by continuous trajectory realization. This design couples reasoning and planning through shared latent computation while maintaining separate tokenizations for language and motion. Training combines supervised reasoning-and-planning objectives, and GRPO is evaluated as an optional post-training extension. Experiments on NAVSIM and Bench2Drive show that RAPDrive outperforms prior non-world-model driving VLA baselines in the supervised/base setting, benefits further from planner-centric post-training, and is supported by ablations on the main architectural choices.