CrewForge: LLM-Forged Crews Measured Against Human-Forged Crews on the Chessboard
Abstract
Multi-Agent LLM Systems (MAS) are typically handcrafted, requiring domain experts to define both agent roles and how agents communicate. Recent work automates role generation using meta-agents, but communication structures often remain fixed or learned, limiting their ability to transfer across tasks. Moreover, evaluations commonly compare automated MAS only against Single-Agent Systems (SAS), making it difficult to determine whether improvements come from multi-agent collaboration or from the automatically generated team itself. We introduce CrewForge, a training-free MAS framework in which a meta-agent can, in a single runtime pass, generate agent roles and specify which other agents each agent can directly communicate with. Given a task, agents deliberate sequentially through direct agent-to-agent messages and a shared space, while the meta-agent supervises the discussion and determines when consensus has been reached. We evaluate CrewForge on chess, a challenging long-horizon environment in which decisions affect future states, game outcomes are objectively determined, and each move can be scored using a chess engine. We evaluate three systems: a SAS, a generated MAS, and a handcrafted MAS with roles informed by literature on chess expertise, allowing us to compare the effects of multi-agent collaboration and automated crew generation. Both the generated and handcrafted MAS win more games than the SAS despite similar per-ply move quality. Measured against the SAS alone, this gain would have been credited to automated team generation, when in fact both MAS share it. A refined deliberation protocol reduces token usage to roughly one-third while maintaining a similar pattern in game outcomes and improving per-ply move quality. Code will be released upon acceptance.