Interaction Value Inference for Multi-Agent Reinforcement Learning via a Hierarchical Agent-Centric World Model
Abstract
In non-stationary multi-agent settings, concurrent policy updates continually reshape inter-agent coordination patterns. In cooperative tasks, an agent’s current action value is closely tied to teammates’ current and future behaviors, making it important to infer how these behaviors evolve and how they influence local action-value estimation. However, existing methods often lack a value-level representation that connects evolving teammate-dependent interactions to local action-value learning. Firstly, we theoretically show that the individual utilities under Centralized Training with Decentralized Execution(CTDE) paradigm struggle to faithfully characterize coordination-dependent action values, making it necessary to introduce a computable proxy for the missing coordination information. Then we propose Hierarchical Inference Agent-Centric World Model (HIA), a framework that incorporates role information into a Transformer-based world model to predict future teammate interactions from each agent’s perspective. The predicted teammate-action rollouts are transformed into auxiliary interaction utilities and integrated into individual value estimates, enabling each agent to capture how teammates’ future actions modulate its current action value. Experiments on StarCraft Multi-Agent Challenge (SMAC) and SMACv2 show that HIA achieves superior cooperative performance over strong baselines.