Dodge-It: Learning Collision-Aware VLA Models for Robotic Manipulation
Abstract
Vision-Language-Action (VLA) models are increasingly used as general-purpose policies for robotic manipulation, but completing the instructed task requires more than reaching the target state. A policy must also avoid nearby objects throughout the rollout, including approach, interaction, transport, and handover. A common approach is to invoke conventional collision avoidance mechanisms, such as motion planners or safety filters, which provide effective geometric safeguards. However, the resulting system may struggle to produce actions that are simultaneously safe and task progressing. When collision avoidance and manipulation progress require misaligned actions, the system may avoid the obstacle while slowing, disrupting, or even failing the task. To address this, we introduce DODGE-IT, a collision-aware learning framework that leverages collision signals as policy-improvement feedback to strengthen the VLA policy's collision-aware manipulation capability. DODGE-IT first designs Task-Relevant Assessment of Collision Events (TRACE) that detects and identifies task-relevant obstacles from RGB-D observations and turns executions into collision outcomes. These outcomes are then incorporated together with task success signals into representative reinforcement learning frameworks for VLA. We validate DODGE-IT on SafeLIBERO and real robot manipulation tasks, showing improved obstacle avoidance while preserving task completion across simulated and physical settings.