Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
Abstract
Recent offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal structure into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time training. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove that decentralized per-agent transport is consistent with joint value improvement and monotonically recovers inter-agent correlation even when the value decomposition is approximate.