Beyond Structural Unification: A Perspective and Roadmap on Unified Multimodal Models
Abstract
Unified Multimodal Models (UMMs) emerge as a promising way to combine understanding and generation within a single architecture. We argue that current approaches mainly achieve structural unification, which can create a Unification Paradox: shared interfaces, token spaces, and parameters do not guarantee an overall capability gain and may degrade individual capabilities. In this position paper, we identify five critical layers in which unification must be rethought: (1) Asymmetric Input Unification, where models must estimate the utility and reliability of each modality rather than treating all modalities equally; (2) Factorized Representation Unification, where semantic abstraction and fine-grained reconstruction should use separate but coordinated capacity; (3) Stateful Reasoning Unification, where models should maintain and revise a shared multimodal belief state across understanding and generation; (4) Synergistic Optimization Unification, where conflicting understanding and generation objectives require coordinated optimization; and (5) Multi-Centric Systemic Evaluation, where benchmarks must evaluate systemic competence, capability preservation, and cross-task robustness rather than isolated outcomes. We argue that progress requires moving beyond structural unification toward functional unification across the full system.