Generalizing Action-Conditioned Latent World Models with Video Model Rewards
Abstract
JEPA-style latent world models offer an efficient alternative to pixel-space video generation by predicting future representations rather than synthesizing future frames. However, current action-conditioned predictors in such latent world models are typically trained on single or narrowly scoped embodied datasets with limited task, object, embodiment, and action-space coverage, and often fail to generalize beyond the training distribution. We argue that this lack of generalist action-conditioned prediction is a key bottleneck for scaling latent world models. In contrast, large controllable video world models have a clear scale-up path through large-scale video pretraining and encode broad visual-dynamics priors, but they are too computationally expensive to serve as online simulators and lack a direct interface to latent action-conditioned prediction. We introduce GeneralistJEPA, a framework for transferring generalist motion priors from video world models to efficient action-conditioned JEPA predictors, establishing a scalable training path for future latent world models. Instead of directly distilling generated futures, whose latents entangle useful action dynamics with source-mismatched content and generation artifacts, GeneralistJEPA factorizes future latent prediction into source-conditioned content continuation and action-induced dynamics. Controllable video world models are used as same-action environment randomizers to build an offline motion-reward bank, providing diverse motion supervision without treating generated content as an absolute target. Specifically, generated rollouts are encoded with a frozen JEPA encoder, and only their patch-level temporal deltas are distilled as video-world-model motion rewards for predicted latent dynamics. This design preserves the efficiency of latent rollout, since video generation and representation encoding are performed offline, while leveraging the broader visual-dynamics priors of video world models. We evaluate GeneralistJEPA on three complementary benchmarks: EgoDex for held-out subtask generalization and action recoverability, BAIR robot pushing for future-frame latent prediction and action sensitivity, and Physion for physical contact readout. Our results show that video-world-model motion rewards improve latent prediction, action sensitivity, action recoverability, and physical contact readout beyond real-only training and naive generated-latent distillation, addressing a key generalization bottleneck on the path toward scaling up action-conditioned latent world models.