Permission to Stop: Monitor Feedback Controls Whether Agents Ask for Help
Shane Caldwell
Abstract
A monitor that blocks an agent's tool call does more than stop an action: its rejection goes back to the agent as the tool output, and what the agent does next depends on the semantic content of that rejection. We study this closed loop in ImpossibleBench, which mutates SWE-bench tasks to make them impossible to solve without cheating. We add a pre-execution monitor that blocks reward hacking tool calls along with a tool for calling on a human to end the session. On 15 tasks chosen because Qwen3-Coder cheated on them unmonitored, monitor plus handoff takes successful cheating from $7/15$ to $0/15$ without hurting performance on matched solvable tasks ($15/15$). We replicate with Claude Sonnet~4 with an 8 task replication goes from $6/8$ to $1/8$. The result we find most interesting is that the wording of the monitor decides whether the agent asks for help at all. A monitor message that specified the violated constraint led to a handoff in $5/11$ trajectories. A generic monitor message of "blocked by policy" never used the handoff tool. Agents under both scoped and generic blocks noticed the inconsistent tests, but only the scoped monitor gave them grounds to treat the task as broken rather than unsolved, which we read as permission to stop. The samples are small and selected and the evidence is behavioral rather than mechanistic, but the practical point stands: the text a monitor returns is part of the control protocol worthy of study. Code, task manifest, and aggregate data: https://anonymous.4open.science/r/permission-to-stop-peer-review-C735/README.md
Chat is not available.
Successful Page Load