FocusVLA: Hijacking Attention to Break Visual Token Pruning in Vision-Language-Action Models
Yanhui Li ⋅ Qianpu Sun ⋅ Qi Zhou ⋅ Chenru Jiang ⋅ Peiyu Zhang ⋅ Dongxia Wang
Abstract
Vision-Language-Action (VLA) models are becoming an important approach for robotic manipulation, but their long visual token sequences make inference expensive. Visual token pruning is a practical way to reduce this cost, with many pruning methods rely on attention scores to decide which tokens to remove. However, we show that this reliance can be exploited: an attacker can hijack attention scores and turn the visual token pruner into a vulnerability. Based on this insight, we propose FocusVLA, a backdoor attack against attention-based pruning. FocusVLA inserts trigger-controlled attention patterns into a few low-impact attention heads. These poisoned heads have little effect on normal manipulation. Without pruning, the poisoned model still behaves normally when the trigger appears. With pruning enabled, clean performance remains intact. But inputs with the trigger shift attention to task-irrelevant peripheral regions. As a result, the pruner removes task-critical tokens, leading to task failure. We evaluate FocusVLA on OpenVLA-OFT and $\pi_{0.5}$ across LIBERO tasks with ten attention-based pruning methods. For example, on OpenVLA-OFT with ADP pruning, the success rate on LIBERO-10 drops from 90.6\% to 31.0\% under triggered inputs. These results reveal a broader risk: attention is widely used as an importance signal, but it can be manipulated and should not be blindly trusted.
Chat is not available.
Successful Page Load