Holistic EvoLution via Intrinsic eXchange for Unified Multimodal Models
Abstract
Unified Multimodal Models (UMMs) integrate visual generation and understanding within a shared parameter space, yet existing post-training typically relies on external annotations and supervised post-training data, and rarely couples these two capabilities for iterative improvement. We propose Holistic EvoLution via Intrinsic eXchange (HELIX) for Unified Multimodal Models, a two-stage cyclic post-training framework that forms a double-helix coupling between generation and understanding: generation produces controllable visual evidence while understanding provides semantic judgments. A target-driven Data Bridge connects the two stages and supplies reliable training signals for both generation and understanding updates. In Stage 1, the understanding branch provides a self-evaluated likelihood gain reward to optimize the generative policy for stronger semantic faithfulness. In Stage 2, the generation branch supplies reconstruction-based visual evidence, and we refine the understanding branch with Reconstruction and Prompt-Alignment rewards. Extensive experiments on Bagel, trained at 512px, demonstrate consistent gains on compositional and knowledge-intensive benchmarks, outperforming the baseline by +5.6 on T2I-CompBench, +7.1 on GenEval, +1.54 on DPG-Bench and +0.02 on WISE when evaluated at 1024px.