Plan and Patch : Diffusion Language Models for Agentic Planning
Abstract
Recent diffusion language model (dLLM) agents have shown competitive performance with autoregressive (AR) agents while offering lower generation latency after task-specific training. Existing dLLM agents, however, largely follow ReAct-style interaction, generating actions incrementally between successive environment observations. While effective for online interaction, this paradigm uses dLLMs primarily as incremental action generators and does not fully exploit their ability to generate and revise structured sequences non-autoregressively. Plan-and-act methods instead separate plan construction from execution, providing a natural setting for leveraging distinctive dLLM capabilities such as parallel generation and infilling. We introduce PLAN-AND-PATCH, a plan-and-act framework that uses dLLMs for both whole-plan generation and localized plan repair. The planner first generates a persistent, program-like plan through parallel unmasking. During execution, feedback is used to localize the region requiring revision, which the dLLM repairs through infilling while preserving the remainder of the plan. We compare DreamReasoner-8B and Qwen3-8B as diffusion and AR planners, respectively. On ALFWorld and TextCraft, after separate task-specific training, diffusion and AR achieve comparable performance in both plan generation and localized repair, while diffusion reduces mean plan-generation latency by 39–46% relative to AR. On Natural Plan, AR performs better at complete-plan generation, whereas diffusion performs better at localized repair. These results provide evidence that explicit planning and localized repair are promising settings for exploiting the native non-autoregressive capabilities of dLLMs in long-horizon agentic tasks.