EnvTrap: Revealing the Environment-Only Attack Surface in Embodied AI via Consequence-Blind Action Execution
Abstract
As AI models evolve from text and vision to physical agents, embodied AI faces a fundamentally different attack surface. Prior attacks on embodied systems have largely focused on semantic or instruction-level manipulations, such as prompts, adversarial images, and action commands; by contrast, we study environment-only perturbations that require no access to the model or instructions. We reveal a new attack surface: the physical environment itself. An adversary who rearranges objects, without modifying instructions or accessing models, can cause hazardous consequences. In this work, we propose EnvTrap, a diagnostic pipeline that constructs paired safe, trap, and null-trap (benign but misleading) embodied scenarios. We demonstrate that environment-only perturbations raise hazardous-action rates to an average of 86.5\% across multiple vision-language-action (VLA) models, with similar vulnerability patterns confirmed in world models. On a consequence-prediction task, model accuracy remains near chance, while human evaluators succeed easily. We further propose a consequence-aware defense that reduces trap trigger rates by an average of 76.3\% across VLA models in simulation and by 61.7\% on physical robots. This vulnerability arises because current embodied models can recognize scene state but often fail to predict action consequences under altered layouts. Our findings establish environment integrity as a prerequisite for safe embodied AI deployment. Our code and data are available at https://anonymous.4open.science/r/Envtrap-1BA4/