CHAIN: Continual Heterogeneous Cooperation with Information Bottleneck for Multi-Agent Reinforcement Learning
Abstract
Efficient heterogeneous Multi-Agent Reinforcement Learning (MARL) in continuously changing environments is key to advancing MARL from simulation to the real world. However, it remains unclear whether current MARL methods can adapt under such dynamic scenarios. In this work, we show that existing methods face significant performance degradation in a novel continual heterogeneous MARL environment we construct, named slippery Multi-Agent Mujoco. We further demonstrate the connection between this phenomenon and representation redundancy as well as insufficient policy heterogenization. To address this issue, we propose Continual Heterogeneous CooperAtion with Information BottleNeck (CHAIN), equipped with Heterogeneous Policy Bottleneck (HPB). In HPB, we extend the traditional information bottleneck into three parts, fitting, compression and heterogenization. Our HPB encourages agents to learn an efficient latent representation to adapt to changed environments, as well as disentangles specializations of agents in state perception, thereby encouraging heterogenization. We evaluate CHAIN on standard MA-Mujoco, our slippery MA-Mujoco and a multi-task continual MARL benchmark, MEAL, with various challenging tasks. The experimental results indicate that our CHAIN outperforms state-of-the-art heterogeneous MARL methods across key continual learning metrics.