Individually Sensible, Collectively Harmful: LLM Agents in Repeated Commons Games
Abstract
LLM agents will increasingly operate in systems where each agent pursues a local objective while the consequences of their actions are shared. We ask whether individually sensible decisions in such settings preserve collective welfare. We introduce Creeping Trap, a repeated commons environment in which three LLM agents choose how much to extract from a shared system. Higher extraction increases an agent's immediate reward but also raises a shared catastrophe risk that harms both the agents and non-acting bystanders. Across nine LLMs and 1460 episodes, agents consistently extract far more than welfare-maximising planners. In the main confirmatory study, 396 of 400 episodes have negative aggregate welfare, and a fresh-seed replication reproduces this result in 217 of 220 episodes. Yet the agents are not behaving arbitrarily: seven of nine models choose extraction levels within 0.09 of an open-loop best response to the behaviour they observe, and median value-space regret is only 1.4%. The result shows that locally sensible behaviour can still produce severe system-level harm. We also find that the surrounding institution matters. Formally payoff-equivalent prompt formulations shift extraction substantially. Making otherwise powerless bystanders more salient reduces extraction across all eight commercial models tested; on Sonnet 4.6 and GPT-5, the direction also holds across three prompt variants. Shorter accountability horizons increase extraction and produce strong end-of-term defection. These results suggest that multi-agent safety requires evaluating not only individual agents, but also the incentives, information, and institutional structures through which their actions combine.