Predicting What Changes: Causal Delta World Models for Risk-Aware LLM Agent Planning
Abstract
Modern LLM agents repeatedly fall into a familiar trap: they execute irreversible actions whose consequences they did not anticipate, such as confirming a non-refundable booking or sending an unrecoverable message. World models are a natural remedy, yet two dominant lines of work both fail this use case. Generative pixel and HTML world models pay a heavy compute cost to produce vivid but often unfaithful continuations of the full observation, while reward and value models only summarise how good an action is and tell the agent nothing about what it will change or whether the change can be undone. We argue that what an agent really needs is a model of what changes, not of what comes next. We introduce Causal Delta World Models (CDWM), which represent each transition as a typed structured delta over an entity, attribute and relation graph, and explicitly factor out reversibility, irreversible cost and prediction confidence. A learnable causal sparsity mask binds each action class to the delta dimensions it can plausibly affect, providing a strong inductive bias for unseen actions and environments. We post-train CDWM with a verifiable composite reward that aligns delta accuracy, action following, calibration error and closed-loop task success. The resulting world model is consumed by a risk-aware closed-loop planner that prunes catastrophic candidates, performs short-horizon model-predictive control, and abstains when confidence is low. On four interactive benchmarks (WebArena, Mind2Web, ALFWorld, ScienceWorld) and two visual control suites (Procgen and Atari 100k), CDWM trained on four NVIDIA A100 (80 GB) GPUs improves task success by 6.0 to 11.7 absolute points over strong agentic and world model baselines, reduces irreversible failures by 47 to 71 percent, lowers token cost per successful trajectory by 27 to 45 percent, and shows substantially better generalisation to unseen websites and rule perturbations. We conclude that the right unit for an agent's world model is the structured delta and not the full observa