DTG-VLA:DynamicTargetGuidancefor Task-CompositionalGeneralizationinRobot Manipulation
Abstract
Vision–Language–Action (VLA) models have shown strong performance on robot manipulation tasks, but they often exhibit weak language grounding and tend to memorize visual and temporal patterns observed during training. As a result, current VLA policies still struggle with task-compositional generalization. For example, a policy trained on individual stages A and B may fail to execute the composed instruction A → B, while a policy trained on A → B may fail when asked to execute the reversed order B → A. We propose DTG-VLA, a dynamic target-guided framework that turns long-horizon instructions into target-conditioned execution stages with fine-grained subtask descriptions. At inference time, DTG-VLA updates both the active target and the current subtask instruction according to task progress, enabling the policy to execute learned interaction stages under new task compositions. Beyond compositional execution,DTG-VLA also improves robustness under visual perturbations by suppressing the task-irrelevant visual context. We evaluate DTG-VLA on real-world robot manipulation and simulation tasks. The results show that DTG-VLA provides a first step toward zero-shot task recombination by enabling non-zero success on unseen task compositions where standard VLA policies and static-guidance variants fail, while also improving robustness under visual perturbations.