ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Abstract
Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction from a short observation window limits temporal context and can bias predictors toward local, low-level extrapolation, making it difficult to capture long-horizon semantics and reducing downstream utility. Vision-language models (VLMs), in contrast, provide strong semantic grounding and general knowledge by reasoning over uniformly sampled observation frames, but they are not ideal as standalone dense predictors due to compute-driven sparse sampling, a language-output bottleneck that compresses fine-grained interaction states into text-oriented representations, and a data-regime mismatch when adapting to small action-conditioned datasets. We propose a VLM-guided JEPA-style latent world modeling framework that combines dense-frame dynamics modeling with long-horizon semantic guidance via a dual-temporal pathway: a dense JEPA branch for fine-grained motion and interaction cues, and a uniformly sampled observation VLM thinker branch with a larger temporal stride for knowledge-rich guidance. To transfer the VLM's progressive reasoning signals effectively, we introduce a hierarchical pyramid representation extraction module that aggregates multi-layer VLM representations into guidance features compatible with latent prediction. Across EgoDex, EgoExo4D, BAIR Robot Pushing, and Physion, ThinkJEPA outperforms diverse latent world model and trajectory prediction baselines across egocentric trajectory prediction, long-horizon rollout, robotic latent prediction, and physical-scene forecasting. These results show that broad visual-semantic guidance from a VLM thinker can benefit JEPA-style latent forecasting.