Nothing to Comply With: A Census of What Developers Specify When They Interrupt a Coding Agent
Ayush Agrawal ⋅ Anish Agrawal
Abstract
When a developer interrupts a running coding agent, what do they actually specify? The answer is in the developer's own words, so measuring it needs no judge. In a uniform sample of 461 such messages, drawn from 7,116 across 1,709 real sessions, 58.6% are directives rather than questions or comments. A tool was executing for 89 of them. Of the directives, 60.5% name an action the developer wants taken and 17.3% forbid anything at all, a presence gap of 43.2 points [27.3, 58.1]. Only 4.4% state a prohibition specific enough for a rule to adjudicate. Two blind human coders agree on the prohibition judgment at $\kappa=0.84$. We reached this question by failing at the obvious one: five attempts to measure whether agents obey each collapsed under a check we ran afterwards. The most transferable is a placebo ladder, with judge, prompt and windows fixed over 117 cases and only the negative control varying. The false-positive rate runs from 0.9% against a message from an unrelated session to 12.0% against the developer's own next queued message. That top rung contains near-duplicates of what was really sent, so we read it as an upper bound; the span is the point. A compliance rate measured this way reports the hardest control its author thought to run. Interruption does not explain the pattern: across those 89 messages, prohibitive language is indistinguishable from length-matched idle turns (18.0% vs 16.0%, $p=0.64$), and we detect no trend over a session. For anyone building monitors or evaluations over agent traces, the prohibition such tools presuppose is usually absent from what the user said. Before asking whether an agent obeyed, check whether it was told anything to obey.
Chat is not available.
Successful Page Load