Reasoning We Can TRUST: Steering Reasoning and Its Effects on VLA Actions
Abstract
As autonomous-driving systems increasingly adopt end-to-end (E2E) foundation models, reasoning-enabled vision-language-action (VLA) models extend this paradigm by generating textual chain-of-thought (CoT) reasoning before predicting trajectories or controls. Although CoT reasoning can improve driving performance, visually ungrounded or behaviorally misaligned reasoning may influence downstream actions and undermine its reliability as an interpretable interface. We introduce Token-level Reward for Utility-Steered Trajectories (TRUST), an offline-trained method that predicts eventual reasoning correctness from partial prefixes and selectively steers unreliable generation while leaving the VLA policy frozen. On Alpamayo 1.5, TRUST accurately monitors reasoning correctness and improves the correctness of steered reasoning from 75.0% to 89.9%. Although this improvement does not materially reduce open-loop trajectory error, TRUST significantly reduces closed-loop trajectory-error metrics and achieves lower error than a compute-matched Best-of-4 baseline. On a baseline-defined challenging subset, TRUST reduces maximum trajectory error by 11.5% and collision rate by 30.4%. We further find that steering changes the frequency of braking intent in the reasoning and that these changes are reflected in subsequent vehicle deceleration. Together, these results identify CoT reasoning as both a source of downstream errors and a viable intervention point for improving closed-loop VLA behavior.