Execution Aware VLAs for Robust Manipulation
Abstract
Most existing VLA approaches assume a fixed and reliable execution interface, where actions predicted by the policy are faithfully realized by the underlying robot dynamics and control stack. In practice, physical robots exhibit tracking error, latency, actuator saturation, compliance effects, calibration inaccuracies, and embodiment-specific dynamics, causing discrepancies between commanded policy actions and the realized robot motion. We refer to this discrepancy as the execution gap. To address this challenge, we propose Execution-Aware Vision-Language-Action (\textbf{EA-VLA}), an approach for adapting pretrained VLA models to varying execution dynamics. EA-VLA augments the robot policy with execution-state feedback, enabling the policy to reason about discrepancies between commanded and executed behavior during action generation. By conditioning on robot proprioception and execution-state information, the proposed approach learns to compensate for the execution gap and tracking error. We demonstrate EA-VLA performance in both simulation and real robot evaluation. In simulation evaluations, we observe EA-VLA maintains performance across perturbation magnitudes two orders of magnitude larger than those tolerated by baseline VLAs. On real robot-experiment we showcase the robustness to underlying controller on Franka and WidowX robot, achieving 11\% higher success on WidowX and around 50\% improvement on throughput on Franka.