Decentralized Diffusion Language Models
Abstract
Discrete diffusion language models (DLMs) generate text by iteratively denoising masked sequences, enabling bidirectional conditioning and parallel token updates. However, current DLM training still follows the centralized paradigm of large-scale language model pre-training, where a monolithic model is optimized across tightly connected GPU nodes with frequent collective communication. We propose Decentralized Diffusion Language Models (D-DLM), a posterior-routed framework for decentralized DLM training. D-DLM partitions the corpus into disjoint data clusters, trains one DLM expert per cluster with no cross-expert gradient or parameter synchronization, and trains a router to estimate the posterior cluster probability of each noised sequence. Our main result shows that, for a data partition, the clean-token posterior defined over the entire corpus decomposes into a mixture of cluster-conditional clean-token posteriors weighted by the cluster association probability. Consequently, the induced discrete probability velocity and categorical reverse process match those of centralized DLM sampling in the realizable full-mixture limit. Our experiments demonstrate that D-DLM reduces cross-node synchronization in an 8-node pre-training setting, improving training throughput by 20% to 24%, and achieving competitive or improved perplexity and downstream accuracy relative to monolithic DLM training.