Nudging the Graph, Not the Agent: A Benchmark for Alignment by Environment Design
Abstract
Agent harnesses are environments: the tools on offer, the ordering of steps and the presence of shortcuts determine not only what an agent can do but what it will do. Practitioners already align agents by editing these environments---removing a tool, forcing a checkpoint, blocking a shortcut---without a formal objective, a solver, or any guarantee that an edit constrains every trajectory the agent might sample. The Kleinberg--Oren model of present-biased planning on task graphs formalises exactly this setting, and the accompanying Minimum-Cost Task Alignment problem asks for the cheapest edit forcing every feasible trajectory to reach the goal and perform a prescribed set of critical actions. The theory is well developed and has never been executed. We release NudgeBench: 18 045 task graphs from four domains---agent tool decompositions, university curricula and scientific workflow traces---annotated with the structural parameters that govern tractability; six parametric instance families whose decision thresholds are placed analytically; exact and heuristic solvers; and an agent-in-the-loop probe that recovers an agent's discount factor exactly rather than by fitting. Real task graphs turn out structurally narrow---treewidth in single digits where prior measurements of network data report hundreds---yet two of seven workflow families remain intractable regardless, because tractability turns on weight diversity as much as on width. The repair an engineer would make by hand matches the optimum on the median real instance and costs up to 2.67x more in the tail, on instances of treewidth two. On general graphs no language-model agent we tested shows present bias, at any capability. But the benchmark's inspection tooling led us to a structural configuration where one model does, reliably and repeatedly---a large step toward the goal beside cheap ways of postponing it---and another does not at all. The behavioural premise therefore holds in an identifiable class of environments rather than universally, while the intervention machinery transfers intact: theory-derived edits move compliance from 0% to 100% across the whole sampling distribution, computed for a discount factor the agent does not have. Code is available at https://anonymous.4open.science/r/NudgeBench-B778