Beyond Training Time, Test-Time Coordination is Essential for Cooperative MARL
Abstract
We argue that cooperative multi-agent reinforcement learning (MARL) should treat test-time coordination as part of the problem definition rather than a deployment artifact. Our focus is on long-horizon, centrally monitored cooperative systems such as warehouse robotics, automated storage-and-retrieval systems, and cloud-resource scheduling, where teams share a reward, operate under shifting objectives, and already run with global telemetry at deployment. The dominant paradigm in this regime, centralized training with decentralized execution (CTDE), uses global information only during training and provides no coordination channel at execution. This is sufficient when agents are weakly coupled or the environment is stationary, but it fails when the team must realign after conditions change. The limitation is structural rather than algorithmic, and more training cannot add a channel that does not exist at execution time. We call this the recoverability gap. To address it, we name the underexplored design space of intermittent test-time coordination as centralized training with mixed execution (CTME), in which a coordinator observes global state and emits high-level signals while agents act locally. CTME is not a particular learning algorithm, but a cooperative MARL problem setting that makes test-time information structure explicit. It specifies team-level information, high-level signals, and update frequency after training, with CTDE recovered as a limit case. We do not argue that CTDE is obsolete. We argue that removing the test-time channel is a modeling choice, not a necessity, and the wrong default when global telemetry is available and teams must recover from drift.