Is Inter-Seed Cross-Play Enough? Assessing the Standard Evaluation Protocol for Zero-Shot Coordination
Abstract
AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot coordination (ZSC) algorithms aim to achieve this by specifying high-level learning rules such that independently engineered agents can coordinate with each other at test time. Rigorous evaluation of ZSC algorithms remains difficult: ideally, multiple independent implementations of each proposed algorithm must be used, reflecting the variation that arises when independent parties interpret and implement the same specification. In practice, however, ZSC algorithms have been evaluated almost exclusively using a single implementation trained across different random seeds, making inter-seed cross-play the de facto standard evaluation protocol. Only a handful of works additionally vary the neural network architecture. This protocol relies on the implicit assumption that inter-seed cross-play is a reasonable proxy for measuring coordination across independently engineered implementations. We test this assumption by introducing cross-implementation cross-play (XIXP), an evaluation scheme that varies implementation details previously shown to affect the performance of multi-agent reinforcement learning (MARL) algorithms, and apply it to Other-Play, a popular ZSC algorithm. For Other-Play with IPPO in Yokai, we find no evidence that the implementation variations we introduce degrade cross-play beyond what inter-seed evaluation already reveals.