$\text{A}^3$-VLA: Automatic Perception-Guided Attention Alignment for Robot Manipulation
Shengzhe Zhang ⋅ Qi Zhang ⋅ Dazhong Shen ⋅ Chao Wang ⋅ Shaopeng Zhai ⋅ Hui Xiong
Abstract
Vision-Language-Action Models (VLAs) enable robots to perform manipulation tasks by leveraging Vision-Language Models (VLMs) to process visual observations and language instructions. However, the action prediction objective itself lacks fine-grained task-specific perceptual supervision, making it prone to relying on task-irrelevant visual shortcuts, leading attention away from target objects and key interaction areas, thereby limiting its operational performance and generalization ability. To address this, we propose Automatic Perception-Guided Attention Alignment for Robot Manipulation ($\text{A}^3$-VLA), a framework that derives task-specific visual guidance from the manipulation process itself. Specifically, $\text{A}^3$-VLA first generates heuristic candidate regions for target objects and the robot from interaction-induced visual changes, and refines them with a visual segmentation model offline. These regions serve as explicit perceptual guidance through two complementary objectives: an attention allocation loss that encourages sufficient focus on task-critical regions when language or robot state tokens query the visual input, and a contrastive learning loss that aligns the representations of candidate target regions with language instructions while separating them from task-irrelevant regions. Extensive experiments on simulated and real-world manipulation tasks demonstrate that $\text{A}^3$-VLA consistently improves multiple VLA backbones, with an average success rate gain of 6.5\% - 25.2\% across benchmarks.
Chat is not available.
Successful Page Load