ActO: Extracting Action Representations from MLLM Embeddings for Video World Models
Abstract
Video world models promise general-purpose interactive simulators, but their training is limited by action supervision: action labels are scarce, fragmented across incompatible specifications, and do not scale with the unlabeled video pretraining corpus. A common workaround jointly trains an action encoder to compress consecutive frames into a latent action and a decoder to reconstruct the next frame from the previous frame and that latent, forcing the latent to capture only new information. This has two limitations: learning the action space from scratch can limit generalization to unseen domains, and the bottleneck trades expressiveness for separation, with tighter settings losing nuance and looser ones risking scene leakage under domain shifts. We argue for a different starting point: rather than learning an action representation from scratch, we adapt a pretrained Multimodal Large Language Model (MLLM), which was trained on large-scale, diverse video-language data and whose embeddings already carry partial action information. This provides a more generalizable foundation for action representation. The remaining challenge is to disentangle action from scene. To this end, we propose a contrastive adaptation framework that reshapes the geometry of the pretrained space without imposing a low-dimensional bottleneck, preserving the expressiveness of the original embeddings. We introduce a unified evaluation protocol that probes action representations along three axes - action expressiveness, scene invariance, and action-conditioned video generation quality - across several robot and game benchmarks in both in-domain and out-of-domain settings. Our representations outperform prior annotation-free latent-action methods on all three axes.