Adaptive Communication Range for Scalable Cooperative Multi-Agent Reinforcement Learning
Abstract
Cooperative multi-agent reinforcement learning (MARL) has long faced scalability challenges due to the exponential growth of the state-action space as the number of agents increases. Existing methods typically enhance scalability by filtering out communications with low-relevance agents. However, such filtering often relies on global state and fixed graph distribution, limiting adaptability in large-scale, communication-constrained environments. In this paper, we propose a scalable MARL method, called Adaptive Communication Range PPO (ACR-PPO), that decomposes the decision-making under communication budget constraints as a sequential process: a communication policy first selects each agent’s communication range within a given budget, followed by a behavior policy that conditions actions on the resulting neighborhood observations. More importantly, we provide a theoretical guarantee of monotonic performance improvement under communication budget constraints. Experiments across diverse scenarios demonstrate that ACR-PPO preserves policy performance while significantly reducing communication costs through adaptive range control.