Before Token Revelation: Linguistic Structure Across Denoising in Diffusion Language Models
Abstract
Recent advances in diffusion-based language models (DLMs) motivate understanding how these models build linguistic structure during denoising. DLMs generate text by iteratively denoising a partially masked sequence, exposing intermediate states that autoregressive models lack. We use these states to ask when, during denoising, two DLMs—Dream-v0-Base-7B and DiffuLLaMA-7B—form interpretable linguistic structure, evaluating them on held-out sentences from the Universal Dependencies (UD) English Web Treebank (EWT), and the German and Japanese treebanks from the Universal Dependencies (UD) Google-Sourced Data (GSD). We report three main findings. (1) Both models contain attention heads that track grammatical relationships even before the participating tokens are revealed. On non-adjacent word pairs (to rule out potential simple positional heads), the selected attention heads identify the correct syntactic governor with held-out accuracies of 45.8–81.1% in Dream and 63.7–93.2% in DiffuLLaMA. When we follow the same word pairs while both remain masked, accuracy increases in 9 of 12 model–relationship combinations and rises by 12.7 percentage points on average over diffusion time. (2) The part-of-speech (POS) and final-token identity become increasingly identifiable throughout diffusion time, including before a masked position is revealed. At 75% denoising progress, the selected later representations reach 78.9% POS accuracy in Dream and 68.4% in DiffuLLaMA, substantially outperforming the implemented controls. Final-token identification also performs best at later depths in both models. (3) Grammatical information is represented by shared, depth-dependent attention patterns rather than separate, universally transferable attention heads. The combined attention entropy (spread of attention) spreads out in earlier layers and becomes more concentrated in deeper layers. Furthermore, English-selected attention heads transfer strongly to German, while weaker on Japanese, showing that these patterns depend heavily on the language and grammatical relationship.