Learning Hierarchical Forward Processes For Discrete Diffusion Language Models
Abstract
Discrete diffusion language models (DLMs) enable parallel generation but usually rely on fixed forward processes, such as masking, uniform noise, or precomputed hierarchies. We investigate whether the hierarchical forward process can instead be learned while preserving continuous-time tractability. We propose \textbf{Learnable Hierarchical Diffusion Language Model (LHDLM)}, which extends a learnable column-stochastic token-cluster map. This many-to-many hierarchy preserves a CTMC with block-conditional transition and closed-form CT-ELBO, while recovering HDLM in the one-hot limit and masked diffusion in the collapsed limit. We identify degenerate hierarchy maps caused by learning the same map used in both forward targets and model-induced cluster predictions, and mitigate them with an information-retention regularizer. Empirically, LHDLM is competitive with discrete diffusion approaches and reveals soft token structures --- cluster-agnostic function tokens and sharp content-token assignments --- that cannot be represented by hard partitions.