Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Kale-ab Tessera ⋅ Andras Szecsenyi ⋅ Cameron Barker ⋅ Alexander Rutherford ⋅ Davide Paglieri ⋅ Aidan Scannell ⋅ Henry Gouk ⋅ Elliot Crowley ⋅ Tim Rocktäschel ⋅ Amos Storkey
Abstract
As language models are increasingly deployed as autonomous agents, they will need to coordinate with others in long-horizon, open-ended interactive tasks. Yet current evaluations rarely test these demands together, focusing instead on short interactions, single-agent open-ended tasks, or highly structured multi-agent settings. We introduce $alem$, a JAX-based, procedurally generated open-ended benchmark for multi-agent coordination built on Craftax-like dynamics. Alem embeds procedurally generated coordination tasks, soft specialisation, communication, and controllable coordination difficulty into a long-horizon survival world with exploration, crafting, trading, and combat. We evaluate $13$ modern LLMs zero-shot within homogeneous teams, with trained MARL agents as reference points. Most LLM agents struggle, averaging only ~6% of maximum reward, although performance varies widely. Gemini-3.1-Pro-High approaches MARL agents trained for one billion steps on Hard coordination ($17.5\%$ vs. $17.6\%$ Coord.\%), Gemma-4-31B-it is the strongest tested open-weight model, and GPT-5.4-High makes strong base progress, while achieving much lower coordination reward. We find that base task competence does not imply coordination competence, communication helps agents share intent, memory and reasoning help with multi-step planning, and heterogeneous teams regress toward average member performance rather than matching their strongest teammate. These results identify coordination as a distinct bottleneck for current LLM agents, separate from single-agent capabilities. Alem makes this bottleneck measurable, providing a controlled setting for developing agents that can communicate, allocate roles, and execute shared plans.
Chat is not available.
Successful Page Load