LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
Abstract
Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs for interleaved cross-modal reasoning is more promising and valuable, e.g., for solving understanding problems that require dense visual thinking, improving visual generation through self-reflection, or modeling visual dynamics of the physical world guided by stepwise action interventions. However, existing UMs necessitate pixel decoding as a bridge due to their disjoint visual representations for understanding and generation, which is both ineffective and computationally cumbersome. In this paper, we introduce LatentUM, a novel unified model that represents all modalities within a shared semantic latent space, eliminating pixel-space mediation for model-internal reasoning over generated visual states. By making generated visual tokens directly interpretable to the model itself, LatentUM enables flexible interleaved cross-modal reasoning and generation. Empirically, LatentUM achieves state-of-the-art performance on the Visual Spatial Planning benchmark, pushes the limits of visual generation through self-reflection and demonstrates world modeling capability by predicting future visual states in the shared semantic latent space.