Masked Generative Pretraining Improves Cross-Dataset Transfer in Pixel-Space Diffusion
Abstract
Pixel-space diffusion models have recently seen a revival in image generation, yet their training remains less efficient than that of their latent-space counterparts. Existing methods reduce this cost by equipping diffusion models with external representations from clean images. In contrast, representation consistency training, a new generative pretraining paradigm, stands out as it eliminates the need of any off-the-shelf semantic encoders. It proposes pre-training the diffusion model to simultaneously capture meaningful visual semantics from clean images while aligning them with data points across various noise levels. Nevertheless, its training depends on contrastive learning with hand-crafted augmentations. This introduces strong semantic prior into the pretraining, limiting model performance when transferred to downstream target distributions that substantially deviate from the pretraining manifold. To address this, we propose Masked Pixel-space Generative Pretraining (MPG), an augmentation-free framework based on masked image modeling. MPG trains the encoder to recover masked image area and aligns those predictions with samples of different noise levels. To measure the performance of pretrained model, we further introduce NC-CKNNA, a metric that quantifies the semantic structure consistency across different noise levels. Under the same architectures and fine-tuning settings, MPG consistently produces better generation performance across seven downstream datasets, while retaining competitive generation quality on ImageNet-256. These results suggest masked generative pretraining as a practical alternative for training pixel diffusion models.