Continuous Latent Diffusion Language Model
Hongcan Guo ⋅ Qinyu Zhao ⋅ Yian Zhao ⋅ Shen Nie ⋅ Rui Zhu ⋅ Qiushan Guo ⋅ Feng Wang ⋅ Tao Yang ⋅ Hengshuang Zhao ⋅ Guoqiang Wei ⋅ Yan Zeng
Abstract
Autoregressive language models have achieved remarkable success in text modeling, but recent work increasingly challenges the fixed left-to-right generation order. However, existing non-autoregressive alternatives such as discrete diffusion, still struggle to jointly deliver efficiency, scalable representation learning, and global semantic modeling. We propose Cola DLM, a hierarchical continuous latent diffusion language model that learns a stable Text VAE, models a global semantic prior in latent space with a block-causal DiT, and decodes text conditionally. Theoretically, under a unified Markov-path formulation, the diffusion operates latent prior transport rather than token-level observation recovery, thereby decoupling global semantic organization from local token realization. Empirically, Cola DLM exhibits strong scaling behavior for text generation across four research questions, eight benchmarks, strictly matched $\sim$2B-parameter AR and LLaDA baselines, and scaling curves up to $\sim$2000 EFLOPs. Meanwhile, we also explore encoding text in the same continuous latent space as images, providing a feasible path toward unified discrete-continuous generative modeling. Our code will be released.
Chat is not available.
Successful Page Load