Learning a Trajectory-Geometric Condition from Reasoning for VLA Planning
Yuguang Yang ⋅ Zhewen Tan ⋅ Canyu Chen ⋅ Cheng Chi ⋅ Chunyang Liu ⋅ Kehua Sheng ⋅ Bo Zhang ⋅ Jinyu Yang ⋅ Linlin Yang ⋅ Baochang Zhang ⋅ Yan Wang ⋅ Xianbin Cao
Abstract
End-to-end driving with vision-language models benefits from multimodal pretraining, but still faces a mismatch between reasoning and trajectory generation. For autoregressive VLA planners, continuous trajectories represented as step-wise normalized text waypoints are strong final outputs because they fit the native token-prediction interface of modern VLMs. However, they remain weak intermediate interfaces for reasoning-conditioned planning. Direct action tokenization provides a more explicit motion interface, but its effectiveness depends heavily on how the action codebook is constructed. We propose *TrajCond-VLA*, which learns a *trajectory-geometric condition* from reasoning for VLA planning. Our key idea is to represent this condition as a differential action trajectory whose $x$/$y$/yaw differentials are discretized separately with a 3D-Brohan codebook. The resulting **DiffAction tokens** preserve local motion geometry while remaining compatible with autoregressive token prediction. Building on this representation, we introduce a **Two-Stage Alignment Training** framework: Stage 1 predicts TrajCond from reasoning, and Stage 2 regresses the final continuous trajectory conditioned on both TrajCond and reasoning. Experiments on NAVSIM, nuScenes, and NeuroNCAP show that the learned trajectory-geometric condition is more effective as an intermediate condition than as a direct output format, and that the resulting pipeline is validated across multiple datasets with consistent gains on the completed comparisons. The nuScenes and NAVSIM navtest discussions are kept aligned to the public benchmark protocols used by Curious-VLA, and the closed-loop reference discussion uses the correct NeuroNCAP protocol.
Chat is not available.
Successful Page Load