Reading and Steering a Vision Language Action Policy in English with Action-Grounded Supervision
Bill Cai ⋅ Minsoo Khang ⋅ Quinn C Han ⋅ Xiaogang Wang ⋅ Yash Shah
Abstract
Vision language action (VLA) policies map images and instructions to robot commands, but their internal representations are difficult to inspect and modify. We build a language interface for a VLA policy whose weights remain fixed. A verbalizer, a separate language model, converts an internal activation into a structured description, and a reconstructor maps the original and edited descriptions back into activation space. Their difference provides a steering vector. We train both modules using descriptions generated from images, task instructions, and the policy's predicted action sequences. Including these actions in the labeling prompt increases mentions of predefined direction words from $6$ of $400$ held-out frames to all $400$. For $\pi_{0.5}$, generated direction words agree with commanded motion on $77.2\%$ of evaluated axes, compared with $54.6\%$ after shuffling frame pairings. Action-conditioned labels also improve scene descriptions, although substantial hallucination remains. Applying direction-word edits throughout simulated episodes steers $\pi_{0.5}$ on three of four tasks and, in a separate control suite, along both tested axes, with larger effects than controls matched for perturbation magnitude. A second policy, GR00T, shows significant steering in one of eight task-reconstructor conditions. Stronger steering reduces task success. These results demonstrate language-based inspection and steering of a frozen VLA and measure its limits in descriptive accuracy, cross-policy transfer, and task preservation.
Chat is not available.
Successful Page Load