Ad-Hoc Teamwork from Human Demonstrations
Abstract
Human players in cooperative games often rely on conventions: reusable patterns of play that coordinate well when used by the whole team, but can fail otherwise. Evaluating agents in this setting is difficult because human demonstration datasets are usually unlabeled: trajectories do not identify which player, group, or convention produced them. As a result, standard proxy-based evaluations can miss important population diversity. In particular, proxy bots produced by regularised self-play around a pooled behavioural cloning policy can remain concentrated around a single dominant mode of play. We introduce a pipeline for recovering and evaluating this hidden convention diversity from unlabeled demonstrations. Our method first learns latent convention structure from the dataset, then selects a compact set of distinct but plausible conventions, and finally reconstructs each selected convention into a high-performing playable proxy using KL-regularised PPO. A key practical challenge is that KL-regularised PPO is sensitive to the KL penalty, which is often tuned using online interaction. We provide an empirical offline calibration procedure for choosing this KL penalty in the PPO+KL reconstruction settings studied here. On Hanabi, we find that the BC-derived proxy suite is highly concentrated, and that the relative ordering of agents changes when evaluation uses the latent proxy suite instead.