Coordination Connectivity: Shared Initialization Shapes the Joint-Policy Landscape in MARL
Abstract
Mode connectivity studies reveal that high-performing neural networks are often connected by low-loss paths in parameter space. We investigate the analogue of this phenomenon in cooperative multi-agent reinforcement learning (MARL), where a solution is not a single model but a team of decentralized policies. We introduce coordination connectivity, a diagnostic evaluating whether linear interpolation between the actor parameters of two trained policy teams preserves high team return. The resulting coordination barrier measures whether two successful teams are separated by a performance cliff along the joint-policy path. We evaluate this diagnostic on the StarCraft Multi-Agent Challenge (SMAC) across five controlled training histories, isolating the effects of shared initialization, continued training, parameter decoupling, and HAPPO-style separate-actor specialization. Across three SMAC maps and four seeds, policy teams derived from a shared MAPPO checkpoint exhibit near-zero interpolation barriers, even after non-shared specialization. In contrast, interpolation paths involving independently trained HAPPO teams exhibit substantially larger barriers. Mixed-team assembly, single-agent hot-swap, clone-all, role-swap, and weight-space probes further indicate that shared-initialized agents specialize into distinct roles while remaining within a low-barrier coordination component. These results suggest that shared initialization implicitly structures the joint-policy space of cooperative MARL, providing a mode-connectivity-inspired lens for analyzing coordination, specialization, and same-source cross-version team compatibility.