Code Comments as an Alignment Surface: Stigmergic Steering of Coding Agents
Jonathan Eicher ⋅ Rafael F Irgolič
Abstract
Many codebases are now generated by Large Language Model (LLM) agents, with subsequent agents deriving a majority of their context from said generated code. Information passed through semantic channels (comments, identifier names, directory structure) paves the way for abstractions that assist in reasoning about the code. Such abstractions can be helpful or harmful; a string can be accurate documentation or a malicious prompt injection depending on the context. In a comparison of 52 agent-generated and 45 human-written repositories, agent codebases are built in an early burst of construction followed by low-churn, append-only maintenance: deletions and renames make up only ${\sim}2\%$ of agent vocabulary turnover (humans: ${\sim}21\%$), so early semantic choices dominate long-term semantics because later agent work rarely removes them. Here we explore the effect of comments on AI coding agents in SWE-bench Verified via comment ablation and adversarial comments. Comment ablation reveals that comments are semantically active but not uniformly aligned: across two models, removal helps one and harms another. We then evolve a set of comments that direct agents to not complete their task. When these comments are injected in place of the originals they reduce task completion by up to ${\sim}28\%$; trajectories in which agents attempt to remove the directive resolve at roughly half the rate of those that leave it in place. Taken together, these results cast the codebase as a stigmergic medium: agents steer other agents through traces deposited in a shared, append-only environment. Directives placed there survive by default, act even when unacknowledged, and are costly to remove; we term this indirect control surface codebase alignment.
Chat is not available.
Successful Page Load