Can Agents Learn from Their Own Interactions? Bootstrapping Strategic Memory from Self-Play
Sana Ayromlou ⋅ Pradyumna Narayana ⋅ Boqi Chen ⋅ Gaurav Kumar ⋅ Srihari Wuntakal
Abstract
Can language agents generate the strategic experience needed to improve future agents without human demonstration data? We investigate this question in multi-turn debate, asking whether external strategic memory induced purely from agent self-play can substitute for memory curated from human interactions. To evaluate both sources under an identical mechanism, we introduce Archetype Retrieval, a framework that extracts structural pairs of argumentative flaws and abstract counter-strategies, projects them into a continuous latent space, and identifies extremal strategic archetypes via Principal Convex Hull Analysis (PCHA). Holding this architecture fixed, we systematically vary only the source of conversational experience across symmetric and asymmetric pairings of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite on the ChangeMyView benchmark. For Gemini 3.6 Flash, memory induced entirely from matched self-play achieves functional parity with human-derived memory in win rate (47.62\% vs. 44.29\%) while significantly accelerating consensus from 15.94 to 15.49 rounds. For Gemini 3.5 Flash-Lite, human memory remains more effective (79.52\% vs. 74.29\%), revealing a capability-dependent induction threshold. Furthermore, cross-manifold analysis demonstrates that effective self-play memory does not simply imitate human discourse: while agents and humans overlap moderately when identifying argumentative flaws ($\sim$55--61\%), their strategic counter-blueprints diverge into largely disjoint spaces ($\sim$35--43\%). Finally, ablating data provenance indicates that memory induced from already-guided interactions provides no consistent advantage over memory induced from raw exploration. Together, these findings suggest that frontier agents can autonomously bootstrap actionable strategic memory along an orthogonal manifold, while highlighting that guided rollouts do not necessarily produce superior downstream learning trajectories.
Chat is not available.
Successful Page Load