ShadowVLA: Test-Time Re-Grounding of Frozen VLA Policies from Self-Generated References
Abstract
Frozen vision--language--action (VLA) policies can perform manipulation tasks under their training layout, but often fail when the same objects are moved by tens of centimeters---even though the instruction, embodiment, and required skill are unchanged. We propose ShadowVLA: a test-time layer that restores execution capability without modifying the policy. We first run the policy in a canonical scene where it is in-distribution, and save a rollout verified by a success predicate as the behavioral reference. In the actual rearranged scene, this reference is decomposed into phases according to the anchor object of the motion, then transported into the rearranged scene via per-phase rigid maps estimated by RGB-D registration, and tracked and executed by a state-synchronized closed-loop controller. The method requires no weight updates, no access to training data, and no additional human demonstrations. On two frozen VLAs and three LIBERO-Goal rearrangements, direct execution succeeds only 3 times out of 59 trials, while re-grounded execution succeeds 51 out of 59. Evaluation is conducted in simulation; the method assumes the existence of a renderable canonical scene.