Block the Heads: Improving the Robustness of VLAs to Spatial Object Perturbation via Attention-Head Blocking
Yunsu Lee ⋅ Kisung Shin ⋅ Suhyung Choi ⋅ Hyundo Lee ⋅ sujin jeon ⋅ Byoung-Tak Zhang
Abstract
The out-of-distribution (OOD) robustness of Vision-Language-Action (VLA) models has recently drawn increasing attention. VLA performance degrades sharply even under small perturbations which constitutes a major limitation for robotic deployment. In this paper, among the OOD problems of VLAs, we focus on spatial object perturbation: when the object's position at inference time differs from its position in the training environment, the model fails to manipulate the object, or the arm even moves toward the training-time position. We first use Delta-probing to test whether information about the object's position is linearly present at each layer of the VLA. We then use Donor patching to test whether each layer actually consumes this information and reflects it in the action. Our analyses show that the cause of failure is not the absence of information but its non-use, and we hypothesize that this arises from competition among a subset of attention heads. Since attention is the operation that moves information between tokens in a transformer block, we narrow down, via head screening, which heads pull the robot toward the old position, and re-run the previously failed tasks closed-loop with that set blocked to test for recovery. We experiment with OpenVLA, OpenVLA-OFT, $\pi_0$, $\pi_{0.5}$, UniVLA, and ACT in the LIBERO environment. As a result, recovery was specific to head selection in only three of the six models---OpenVLA-OFT improves from $2.5\to 14.2$\%.
Chat is not available.
Successful Page Load