Joint Adaptive Neighborhood Constraint for Offline Multi-Agent Reinforcement Learning
Abstract
Offline reinforcement learning suffers from distributional shift and extrapolation errors. These issues are particularly severe in multi-agent settings. As the number of agents increases, the joint action space grows exponentially while agent behaviors become highly coupled. Consequently, even if individual actions remain within the data distribution, their combination may still result in out-of-distribution (OOD) joint actions. Directly extending single-agent constraints to multi-agent settings is difficult, as they fail to effectively constrain joint actions. This often results in poor coordination or excessive conservatism. To bridge this gap, we propose the joint adaptive neighborhood constraint (JANC), which explicitly constructs controlled neighborhoods in the joint action space to suppress extrapolation while preserving reliable generalization. Moreover, JANC adaptively scales neighborhood radii based on joint advantages to align generalization with value structures. In practice, our method first performs adaptive neighborhood Q learning to explore high-value behaviors and then conducts value and policy learning based on the refined actions. Experiments on offline benchmarks, including multi-agent Mujoco and StarCraft II, show that JANC outperforms state-of-the-art methods on the majority of tasks in complex cooperative scenarios. We will release our source code to the public.