A Latent World-Action Model with Jointly Aligned Reasoning
Abstract
Despite recent progress, existing general robot policies, particularly Vision-Language-Action models, are still primarily trained to map observations directly to actions using sparse action supervision. As a result, they often learn shallow observation-to-action correlations instead of deeper representations of world dynamics, including object motion, physical interaction, and long-horizon task progression. Recent world-action models attempt to address this limitation through future-frame prediction, but pixel-space rollout is computationally expensive and poorly aligned with the abstractions required for control, forcing the model to reconstruct visual details irrelevant to action. We present Being-H0.7, a latent world-action model that introduces future-aware reasoning into VLA-style policies without generating future frames. Our model inserts learnable latent queries between perception and action, forming an explicit reasoning interface for action generation. To train this interface, we adopt a future-informed dual-branch design: a deployable prior branch infers latent states from the current context, while a training-only posterior branch replaces latent queries with embeddings from future frames. Joint alignment between the two branches enables the prior branch to learn predictive, action-relevant latent structure directly from current observations. Under a controlled training setup for comprehensive validation, Being-H0.7 achieves state-of-the-art or competitive performance across diverse benchmarks, while further demonstrating strong deployability on real-robot manipulation tasks.