Coordinating Hundreds of RL Agents through Scalable Inference-Time Search
Abstract
Multi-Agent Reinforcement Learning (MARL) is a powerful approach to large-scale systems control, such as power grids and traffic networks, where dozens to hundreds of agents must coordinate effectively. Furthermore, such systems often operate in digital or simulated settings, which allows for inference-time search to significantly improve upon zero-shot performance. A current leading approach in this setting is COMPASS, which builds a continuous family of policies conditioned on latent vectors that can be efficiently searched at test time, but requires all agents to share a single latent vector, limiting the space of reachable joint policies. We introduce ATLAS, a simple extension in which each agent conditions on its own latent vector while still optimising a joint objective. This expands the searchable policy space at negligible computational cost and requires no additional training. ATLAS consistently sets a new state-of-the-art on a benchmark of 14 scenarios across three environments, scaling to teams of up to 200 agents, widely out of distribution. Crucially, its lead sharpens with team size and compute, reaching up to a 45% performance gain at scale.