A Multi-Agent Research Lab for GLEE
Abstract
GLEE evaluates agents in bargaining, negotiation, and persuasion. We entered four deterministic players, each revised between games by a persistent language-model researcher agent. A coordinator agent connected the researcher agents through shared records and communication, forming a research mesh. Within human-supplied infrastructure, the agents built diagnostic tools, replay tests, comparison scripts, and channels for findings and retractions. We examine this organization through its code, discussions, and 247,057 rated completions over 111 hours. Bargaining openings converged through code sharing, while all four persuasion trajectories ended above their starting ratings. The lab questioned its own analyses and corrected claims, yet its checks did not consistently govern policy allocation or evaluate complete player programs. The case exposes a gap between establishing scientific practices and making decisions conditional on their results.