SamaDICE: Safe Multi-Agent Reinforcement Learning with Stationary Distribution Correction Estimation
Abstract
Offline safe multi-agent reinforcement learning (MARL) is challenging: agents must learn coordinated behaviors from static datasets while strictly satisfying safety constraints. This challenge is amplified by distribution shift, where joint actions that appear safe under the behavior distribution may become unsafe under policy-induced deviations, with errors compounding across agents. We propose SamaDICE, a principled framework for safe offline MARL that integrates stationary distribution correction with scalable value decomposition. Our approach formulates safe policy learning as a constrained convex optimization problem over stationary distributions of joint state-action pairs, with safety constraints imposed as upper bounds on expected cumulative costs. Safety is enforced through a Lagrangian dual formulation, enabling distribution-corrected cost estimation directly from offline data. To address the combinatorial complexity of multi-agent systems, we introduce a factored centralized-training-decentralized-execution parameterization that represents the global dual variable via a monotonic mixing network over agent-local potentials, yielding a tractable surrogate Bellman residual for scalable optimization. We instantiate SamaDICE with a multi-agent decision transformer and evaluate it on diverse benchmark environments. Our method outperforms the strongest baseline in aggregate return, with the advantage widening as the cost budget tightens, while maintaining 100% empirical episode-level constraint satisfaction across all evaluated tasks.