Riemannian Cooperation Dynamics for Stable Multi-Agent Adaptation
Jagannatham Jahnavi ⋅ Suprith S
Abstract
Cooperative multi-agent reinforcement learning (MARL) systems increasingly operate in dynamic environments where agents must continuously adapt under evolving tasks, reward functions, and interaction dynamics. However, simultaneous policy updates in non-stationary settings often induce oscillatory learning behavior, unstable equilibria, and coordination collapse. Existing MARL methods primarily optimize policies in Euclidean parameter spaces, which inadequately capture the intrinsic statistical geometry governing cooperative policy evolution. In this work, we introduce a geometric framework for stable cooperative adaptation by modeling multi agent policy evolution on a Fisher-Riemannian statistical manifold induced by the Fisher Information Metric. Rather than interpreting policy learning purely as parameter optimization, we view cooperative adaptation as smooth movement along geometry-aware trajectories in policy space. This perspective provides a principled interpretation of coordination stability through local statistical structure and policy manifold dynamics. Building on this formulation, we propose a lightweight Fisher-geometry regularization mechanism that stabilizes adaptation by constraining temporal drift in empirical Fisher structure across consecutive policy updates. Concretely, the method penalizes abrupt changes in local policy geometry using Frobenius distance regularization between empirical Fisher matrices, encouraging smoother and more stable cooperative evolution under environmental perturbations. The resulting optimization objective augments standard MARL learning with a geometry-stability term: $$ L = L_{\mathrm{RL}} + \lambda \sum_{i=1}^{N} \|\hat{g}_{i,t} - \hat{g}_{i,t-1}\|_F $$ where $L_{\mathrm{RL}}$ denotes the base reinforcement learning objective and the additional regularization term constrains instability in local policy geometry during training. Unlike computationally expensive second-order geometric optimization approaches, our framework relies on tractable empirical Fisher approximations and integrates efficiently into existing MARL pipelines with minimal overhead. We evaluate the proposed framework on non-stationary cooperative navigation environments derived from the Multi-Agent Particle Environment (MPE), incorporating dynamic reward perturbations, changing task assignments, and evolving coordination objectives during training. The proposed regularization is integrated into both MADDPG and MAPPO frameworks and compared against standard as well as natural-gradient variants. Preliminary experiments are designed to evaluate coordination stability, return variance, inter-agent policy divergence, and robustness under environmental perturbations, with the goal of assessing whether geometryaware regularization improves cooperative adaptation in non-stationary settings. Beyond empirical performance, this work highlights the broader potential of information geometry as a principled foundation for reliable cooperative AI systems. By incorporating geometric structure directly into policy adaptation dynamics, the proposed framework offers a scalable and interpretable approach for stabilizing multi-agent learning in evolving environments, opening new directions at the intersection of information geometry, reinforcement learning, and adaptive cooperative intelligence.
Chat is not available.
Successful Page Load