Addressing Exogenous Variability in Cooperative Multi-Agent Reinforcement Learning
Abstract
Cooperative multi-agent reinforcement learning (MARL) often fails under exogenous variability such as shifts in opponents or environment regimes that cannot be controlled by the team but alter the transition dynamics. We formalize this challenge as Exogenous Dec-POMDP (ED-POMDP), which decomposes the global state into endogenous variables controllable by the team and exogenous variables beyond its control. This formulation exposes delayed influence: While actions do not directly determine exogenous transitions, they can shape them through induced changes in endogenous state. Based on this, we propose LEICA, a CTDE-compatible algorithm that learns history-conditioned endogenous and exogenous context representations and shapes policy updates using influence-weighted intrinsic rewards. Across SMAX benchmarks with opponent strategy shifts, LEICA consistently improves both training performance and generalization to unseen opponents' strategies over existing baselines, supporting the usefulness of influence-weighted shaping under exogenous train-test regime shifts.