Where a Conflict Arrives: What Decides Which Instruction a Coding Agent Follows
Abstract
A coding agent reads instructions from at least seven places: the system prompt, an AGENTS.md file, the README, code comments, sibling modules that demonstrate a style, and the user's first and latest messages. Documentation asserts a precedence order among them. We put one style instruction in each place and set all 21 pairs in conflict, scoring every artifact deterministically from the code it produced, never with an LLM judge. Across 10,835 controlled trials, 6 models and 3 families, no order holds throughout. The endpoints are stable, user messages near the top and repository files at the bottom. The contested middle (the system prompt against AGENTS.md) reorders across families, inside one family, and inside a single model when only an inference setting changes. Where the hierarchy fails, position decides. Holding content fixed and swapping two instructions' order within the conversation reverses which one wins in all 6 models; the same swap inside the system prompt is inverted, weak, or undetected. Agents also sometimes satisfy both. Told to raise on invalid input, then to return None, one agent did both. It returned None only below absolute zero, a split neither instruction mentioned. Only 1 of the 21 pairs does this at a substantial rate: the two conversational turns (0.217, against 0.003 for two repository files). One family never does it (0/960). Both results fall on the same boundary, between the conversation and everything else.