DDACOM: Dual-Decoupled Adaptive Communication for Multi-Agent Reinforcement Learning under Dynamic Networks
Abstract
Communication enhances cooperative multi-agent reinforcement learning (MARL) under partial observability. In real networks, however, each agent's available bandwidth varies over time due to changing link quality, interference, and congestion. Most existing MARL communication policies decide whether to transmit based on message importance or novelty, or are trained under fixed bandwidth conditions. As a result, they cannot adapt sending frequency to bandwidth variations, leading to packet loss when bandwidth decreases or missed information when bandwidth recovers. To address this, we propose Dual-Decoupled Adaptive Communication (DDACOM)---a framework that decouples message sending and message generation from the action policy. The message-sending module ranks each message by its deviation from teammates' information and the sender's history, and uses a dynamic token bucket to convert the probe-estimated safe rate into a real-time token budget, enabling the communication frequency to track the safe rate online. The message-generation module is trained independently with receiver-side attention feedback and counterfactual marginal value, ensuring each transmitted message meaningfully contributes to cooperation. On SMAC and MPE, DDACOM achieves leading cooperative performance among baseline methods. Under dynamic network conditions in ns-3 simulation, DDACOM sustains over (95\%) safe-rate utilization and packet loss below (0.3\%) while reducing communication frequency to less than half of full broadcast, confirming its adaptability to runtime network variation.