Boosting Multiagent Reinforcement Learning at High Replay Ratios with Ensemble Reset
Abstract
Increasing the replay ratio, where an agent's network is updated multiple times per environment interaction, is an effective strategy for improving sample efficiency in reinforcement learning. However, its effects on the network capacity of multiagent reinforcement learning (MARL) are not yet well investigated. In this paper, we show that high replay ratios induce a severe dormant neuron problem in the centralized global Q-network of MARL, where a large fraction of neurons become inactive, thereby reducing network capacity and destabilizing learning. To address this problem, we propose Ensemble Reset (EnSet) to stabilize MARL training at high replay ratios. First, guided by theoretical analysis, EnSet employs an ensemble of global Q-value networks with periodic resets to mitigate neuron dormancy during frequent updates. Second, EnSet diversifies replay experience by exploiting multiagent translation invariance as an inductive bias in the global Q-value function to prevent overfitting. Extensive experiments in SMAC, MPE, and SMACv2 environments demonstrate that EnSet consistently improves various MARL algorithms at high replay ratios with fewer environment steps.