OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
Abstract
Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes latency prohibitive for real-time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We argue that this is because purely linguistic latent representations compress a symbolic abstraction of the world rather than the causal dynamics that govern driving. We present OneVL (One-step latent reasoning and planning with Vision-Language explanations), a unified VLA and world model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. It comprises a language decoder that reconstructs text CoT and a visual world model decoder that predicts future-frame tokens, forcing the latent space to internalize causal scene dynamics. A three-stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives. During inference, the decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching answer-only prediction speed. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering state-of-the-art accuracy at answer-only latency. Code will be publicly available.