DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
Xu Huang ⋅ Ye Huang ⋅ Zijun Liao ⋅ Yuwei Niu ⋅ Xiaojie Li ⋅ Menghan Zhou ⋅ De Wen Soh ⋅ Xiaotong Li ⋅ Daquan Zhou
Abstract
High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increase the learning diffucity of diffusion training, resulting slow model convergence. Recent representation autoencoders speed-up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. % Specifically, on ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with \textbf{29.79} PSNR and \textbf{3.05} gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5\% and 59.7\% on PSNR and gFID respectively, maintaining comparable throughput and faster diffusion model training convergence.
Chat is not available.
Successful Page Load