Representation-Space MMD for Few-step Continuous Diffusion Language Models
Abstract
We introduce a kernel-based distribution matching approach for few-step continuous diffusion language models (CDLMs). Starting from a pretrained CDLM, we train a one-step generator to directly match the real and generated distributions by minimizing Maximum Mean Discrepancy (MMD) in a representation space induced by the frozen teacher. Our method requires neither paired teacher trajectories nor a jointly trained discriminator or score estimator. At inference time, we leverage self-conditioning to iteratively refine the one-step predictions, enabling effective generation across different inference budgets. Applied to the recent Embedded Language Flows (ELF) model, our approach outperforms prior few-step CDLM methods on both unconditional OpenWebText generation and conditional mathematical reasoning on GSM8K.