Learning Robot Behavior Priors from Non-Robotic Synthetic Data
Abstract
Real robot data is often treated as the bottleneck for training robot learning models, including embodied foundation models. Progress toward zero-shot physical AI is largely driven by scaling embodied data across tasks and morphologies, yet physical experience data is costly to collect and platform-dependent. In this paper, we investigate whether any of the structure that such models require can instead be acquired from data originating from systems without embodiment. To this end, we test if inexpensive, procedurally generated synthetic data can yield useful priors before a model observes any real-world behavior. We study this through robot action prediction, a natural pretraining objective and a direct test of whether structure learned from arbitrary synthetic tasks transfers to embodied settings. We introduce three data generation pipelines that model fundamental abstractions of robotic control, and evaluate how training with these pipelines compares to the generic causal task generation proposed by recent work. We demonstrate that a model trained purely on this non-robotic data attains non-trivial zero-shot prediction performance on unseen physical robots via an in-context learning mechanism, and our specialized control priors notably improve upon the generic causal prior. We also compare against a compute-matched baseline trained on real robot data. The baseline outperforms our model on the datasets it was trained on, but it fails to transfer to robotic datasets held out from its training, where our model performs better—a collapse that may reflect our limited model scale. These findings suggest that procedurally generated, mathematical abstractions can yield useful behavioral priors, offering a potential complement to expensive embodied data collection.