Trace–Answer Compatibility Emerges at Depth and Mediates Prediction in Diffusion Language Models
Ashutosh Ojha ⋅ Pranath Reddy ⋅ Sergei Gleyzer
Abstract
Diffusion language models expose latent correctness signals in their hidden states, but it is unclear what those signals represent or whether the model uses them. Studying LLaDA-8B and Dream-7B with autoregressive controls (Qwen3-8B and Llama-3.1-8B), we identify a concrete representational object: late hidden states encode the compatibility between a generated trace and its final answer. A $2\times2$ intervention crosses the non-answer trace from correct- and wrong-answer generations with correct and wrong final answers, isolating a trace--answer interaction. Without using correctness labels, compatibility is linearly decodable from deep states at AUC 0.919 in LLaDA and 0.897 in Dream and transfers across equivalent answer formats, while the shallow geometry is weak and poorly aligned. Crucially, this representation is causally predictive: re-masking the answer and patching deep answer-slot activations between correct- and wrong-source traces mediates 66\% and 80\% of the trace-conditioned answer-likelihood gap, with matched pre-answer controls that are null or small. The effect is dose-dependent, requires a question-matched donor, and is sensitive to controlled numerical and ordering corruptions of the trace. Objective-matched AR controls show that compatibility is not unique to diffusion models; in the diffusion models studied, however, it forms later and becomes more causally concentrated at masked-answer slots.
Chat is not available.
Successful Page Load