X-EPA: Scalable Latent World Models with Natural Language Actions
Junha Song ⋅ Byeongho Heo ⋅ Geonmo Gu ⋅ Dongyoon Han ⋅ Sangdoo Yun
Abstract
Pixel and latent world models have developed along separate paths. Pixel world models condition on language and learn from web-scale video, which makes them general, whereas latent world models plan with numerical actions on robot and simulation data, which makes them efficient but domain-specific. We ask whether a latent world model can inherit the generality of language conditioning while retaining the efficiency of latent prediction. We present X-EPA, a latent world model that translates the numerical actions of simulation and real-robot datasets into language and trains a language-model predictor to regress future latent states from these language actions. Video-caption and image-editing transitions supplement the training data so that the predictor grounds language in visual change rather than in recurring patterns. Evaluated on six simulation environments, one real-robot benchmark, and one zero-shot robot benchmark, X-EPA achieves 12\% higher average planning success than state-of-the-art latent world models trained separately for each domain. It also plans with actions expressed solely in natural language. Finally, the shared interface allows a direct comparison of latent and pixel world models under matched conditions: X-EPA matches the pixel model with $6\times$ less training, plans $13.9\times$ faster, and surpasses it by 24\%, favoring latent prediction in the long-standing debate.
Chat is not available.
Successful Page Load