Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
Itamar Trainin ⋅ Shauli Ravfogel ⋅ Omri Abend ⋅ Amir Feder
Abstract
Discrete diffusion models are often said to outperform autoregressive models on planning tasks, but the specific mechanisms enabling this are not yet well understood. To explain why, we abstract the complexity of studying ``planning'' into graph traversal problems, which offer a minimal setup for evaluating language models' planning and lookahead capabilities. We find that non-autoregressive language models are able to leverage an inherent directional asymmetry in lookahead planning tasks to outperform the training efficiency of their autoregressive counterparts. Through a mechanistic study of the training and inference dynamics of autoregressive and non-autoregressive models, we discover that while forward traversal through junctions in the graph requires complex sequential planning, the reverse path remains deterministic. Consequently, non-autoregressive models naturally develop a fundamentally different solution strategy. Effectively, they learn a reverse-decoding pattern which reduces an $\ell^{th}$-order lookahead problem to a sequence of simple $1^{st}$-order transitions. Latent-space analysis further confirms these divergent strategies by demonstrating the distinct internal representations formed during generation. Ultimately, these findings highlight the advantages of non-autoregressive modeling while clarifying the inherent planning limitations induced by autoregressive architectures.
Chat is not available.
Successful Page Load