FreeAct: Demonstration-Free Robot Adaptation via Action-Grounded Generated Videos
Abstract
Adapting pretrained Vision-Language-Action (VLA) policies to new deployment environments typically requires collecting expert demonstrations in each target domain, which is costly and difficult to scale. In this paper, we present FreeAct, a framework for demonstration-free robot adaptation that converts generated videos into physically grounded supervision. However, a key challenge is that generated manipulation videos, while visually plausible, often contain embodiment artifacts and kinematic inconsistencies that make direct inverse-dynamics labeling unreliable. FreeAct addresses this challenge by learning a discrete latent action space from large-scale multi-lab robot videos and aligning it with camera-frame end-effector SE(3) motion and gripper-state changes. The resulting latent action tokenizer captures task-relevant motion while reducing reliance on embodiment-specific visual artifacts, enabling generated target-domain videos to be pseudo-labeled with reliable latent action tokens. These tokens are first injected into VLA pretraining and then used, together with source-domain action supervision, to adapt the policy to unseen environments. Empirically, FreeAct improves success from 2\% to 57\% under target-domain shift using generated videos alone, and further to 74\% through MLLM-guided self-evolution on the policy's own rollouts, closing the adaptation loop without expert demonstrations. Further analysis shows that the learned latent action space forms a kinematically structured and transferable representation, establishing generated videos as a scalable source of supervision for robot manipulation.