VUM: Visual Unified Models for Image Generation and Perception
Abstract
Unifying visual perception and generation poses a fundamental dichotomy: perception extracts semantics but discards details, while generation synthesizes fine-grained structures. Thus, existing works rely on decoupled architectures, as forcing these opposing information flows into one network inevitably triggers task interference and representation collapse. To break this bottleneck, we introduce the Visual Unified Model (VUM), a native framework that unifies both capabilities within a single architecture. Our key insight is that visual perception and generation can be unified by re-envisioning Masked Image Modeling and Diffusion Processes as complementary facets of data recovery. By establishing a shared degradation space that integrates both masking and noising, VUM optimizes a joint signal recovery objective: simultaneously predicting masked patches and denoising corrupted signals. This objective enforces the dual internalization of semantic abstraction and generative synthesis. Using a shared, unified backbone, VUM delivers highly competitive performance across a diverse spectrum of tasks, encompassing image generation, visual recognition and dense prediction tasks. Codes and pretrained checkpoints will be made publicly available.