ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
Abstract
Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety-critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow-VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR-induced modes and learn VLM-conditioned residual distributions over these modes. An autoregressive generator (\textbf{Chain}) produces a discrete set of causal trajectory modes, followed by a diffusion-based refiner (\textbf{Flow}) that operates in residual space to perform mode-conditioned correction while preserving causal structure. A key insight is that vision-language models are more effective as semantic controllers for refinement rather than direct trajectory generators. By conditioning the diffusion process on VLM hidden states, scene-level reasoning guides fine-grained trajectory adjustments within each mode. This formulation enables robust planning in ambiguous and long-tail scenarios. ChainFlow-VLA achieves a state-of-the-art score of 94.8 on the NAVSIM v1 leaderboard, matching human-level performance (94.8).