APEX: Sustaining Exploration in Self-Evolving LLM Agents
Abstract
Self-evolving LLM agents adapt across episodes by retaining experience in external memory while keeping model weights fixed. Yet accumulated experience can also restrict future exploration: agents repeatedly exploit familiar routines and miss better strategies. We study this failure mode, exploration collapse, and propose Autonomous Policy Exploration (APEX), a framework that makes both acquired knowledge and untried alternatives explicit. APEX maintains a strategy map of milestones and prerequisite dependencies. Fork Discovery expands this map using alternatives grounded in interaction history, while uncertainty-aware Policy Selection chooses which milestone to attempt next. Map refinement and return propagation consolidate experience between episodes. On nine Jericho games, APEX achieves the highest reported final-five-episode mean among the evaluated methods. On WebArena, it achieves 55.9% final-three-episode success, compared with 50.5% for the strongest baseline. Ablations support the value of map structure, frontier expansion, and uncertainty-aware selection. These results identify exploration maintenance as a practical component of memory-based agent adaptation; they do not establish resistance to forgetting under changing task distributions.