Learning Large-Scale Competitive Team Behaviors with Mean-Field Interactions
Bhavini Jeloka ⋅ Yue Guan ⋅ Panagiotis Tsiotras
Abstract
While multi-agent reinforcement learning (MARL) has shown strong empirical performance, existing methods struggle with large number of agents, due to the combinatorial growth of joint interactions. Mean-field (MF) approximations address this by replacing pairwise interactions with interactions against population distributions, yielding tractable large-population policies. However, existing MF formulations focus on fully cooperative or purely competitive settings and fail to capture the mixed cooperative–competitive structure of team-based games. We extend mean-field learning to _zero-sum team games_, where agents cooperate within teams and compete at the team level. We show that such games admit $\epsilon$-optimal decentralized policies that depend only on local states and population distributions. Building on this structure, we propose MF-MAPPO, a scalable algorithm with a shared actor and a minimally informed critic per team. MF-MAPPO is trained directly in finite-population simulators rather than using mean-field oracles, thereby enabling deployment to realistic scenarios with thousands of agents. We further extend MF-MAPPO to partially observable settings via a simple gradient-regularized training scheme. Experiments on large-scale benchmarks in our simulation platform $\texttt{MFEnv}$, including population games with analytical solutions and high-dimensional battlefield scenarios, demonstrate that MF-MAPPO outperforms existing MARL baselines and yields rich, heterogeneous behaviors.
Chat is not available.
Successful Page Load