LLM Agents Self-Improve in a Bluffing Game by Writing Notes
Abstract
Large language models can improve from experience without changing their weights if they are allowed to write down what they learn. We study this idea in Perudo, an imperfect-information bluffing game, where frozen LLM agents maintain a persistent notebook of free-form English notes across games. Long-term memory produces large gains: GLM-5.2 and Gemini 3.7 Flash defeat otherwise identical no-memory versions of themselves 44-6 and 47-3, respectively. More strikingly, a 1,459-entry notebook written by GLM-5.2 during self-play transfers beyond the setting in which it was created. Without memory, GLM loses 16-34 to Gemini; given the frozen notebook, the same model reverses the matchup and wins 32-18. The same GLM-written notebook also improves Gemini against another Gemini agent, despite being written by a different model against a different opponent. A size-matched control of the notebook does not reproduce these gains. The resulting notes are readable and insightful but imperfect: they contain general strategies, opponent models, self-directed guidance, and episodic records alongside contradictions and incorrect advice. These results provide a concrete example of non-parametric self-improvement in which an LLM bootstraps useful knowledge from self-play, stores it in plain English, and produces an artifact that remains useful beyond both the model and opponent that generated it.