CoffeeBench: A Benchmark for Long-Horizon Strategic Decision-Making in Multi-Agent Economies
Abstract
As LLM agents advance beyond short-horizon tasks, evaluating their ability to perform long-horizon strategic decision-making in realistic multi-agent settings remains a key challenge. Existing agent benchmarks typically focus on isolated task completion or interactions with rule-based or fixed-policy counterparts, limiting their ability to capture strategic adaptation and emergent behaviors arising from sustained multi-agent interactions. We introduce \textbf{CoffeeBench}, a benchmark for evaluating LLM agents in a dynamic multi-agent economic environment consisting of farmers, roasters, and retailers forming a multi-stage supply chain. The environment consists of autonomous firms interacting over a multi-month horizon with evolving supply, demand, and pricing, creating sustained competition and supply-chain dynamics. In each evaluation, the evaluated model controls one firm and must make sequential decisions on procurement, production, pricing, and negotiation to maximize cumulative net income. CoffeeBench enables the study of emergent behaviors arising from repeated strategic interactions, including negotiation dynamics, pricing discipline, and coordination failures. We evaluate several frontier and pareto-frontier LLMs on CoffeeBench to analyze their economic performance and strategic behavior. Most evaluated models achieve positive net income, with GPT-5.5 achieving the highest mean net income. In contrast, Claude~Haiku~4.5 frequently yields negative net income, exhibiting an idle-drift failure mode in which agents collapse into inactivity despite coherent plans. We further find that stronger models proactively negotiate with counterparties through frequent messaging, whereas weaker models trade reactively with limited communication. CoffeeBench provides a controllable testbed for studying strategic interaction, coordination, and long-horizon behavior in multi-agent economies.