Learning Not to Act: Suppression Training for VLAs Under Invalid Instructions
Abstract
Vision-Language-Action models (VLAs) may take unexpected actions when presented with invalid or unexecutable language instructions, potentially causing harm to the environment or the agent itself. We propose Invalid Instruction Suppression Training (IIST), a training strategy that regularizes VLA behavior in response to invalid instructions. Across eight MetaWorld tasks per task, IIST reduces the mean end-effector path length by 91.9% for seen invalid instructions and 89.2% for held-out invalid instructions. It also suppresses unintended execution of tasks from the training data: the false-execution rate, defined as the rate of completing such tasks under invalid instructions, drops from 12.5% to 1.25% for seen invalid instructions and from 15.0% to 1.25% for held-out invalid instructions. Moreover, our experiments show that IIST preserves VLA performance on valid instructions.