Fixing What You Cannot Find: The Asymmetric Exploration Shortfall of LLM Self-Play
Abstract
Autonomous self-play is widely used to drive self-improving reasoning in language models, yet how thoroughly it explores the space of human argumentation remains an open question. We investigate whether self-play explores problem-critique and solution-generation spaces symmetrically by measuring the high-dimensional coverage of debates on ChangeMyView against held-out human discussions under matched sample sizes. By separating arguments into the reasoning flaws exposed and the constructive blueprints proposed, we uncover a pronounced structural asymmetry: self-play covers human blueprint strategies relatively well (a 2.8\% to 4.0\% deficit), but suffers roughly twice the coverage shortfall when exploring reasoning flaws (5.6\% to 7.9\% deficit, p<0.001). This gap is not caused by models omitting broad argumentative categories—macro-level archetype distributions are statistically indistinguishable from humans (0/16 and 0/24 tests survive FDR correction)—nor by broken deductive transitions, as conditional blueprint entropy matches human baselines to within 0.02 bits. Instead, self-play exhibits a localized geometric contraction specifically in flaw discovery: models know how to blueprint arguments once framed, but fail to reach the diverse frontier of human flaws. Furthermore, while pooling outputs across diverse model pairings adds genuine geometric diversity, naively aggregating them under a fixed sample budget triggers a severe selection penalty: degenerate pairings act as capacity sinks that erase diversity gains (leaving a naive four-way mixture 8.5 percentage points behind a curated subset on flaws). Our findings establish that self-play's core bottleneck is problem-discovery rather than blueprint generation, and that expanding coverage requires geometry-aware curation rather than unweighted model ensembling.