Curated Learning for Self Improving Agents
Abstract
Large Language Models (LLMs) deal with multi-turn agentic workloads more often than single turn QA. Existing reflective prompt optimizers such as GEPA learn from failed rollouts, but agentic failures can come from prompt weaknesses, harness limitations, stochastic execution, or noisy judges. We argue that such failures may not provide useful learning signal, and may even be harmful, and it is essential to curate which tasks provided to the prompt optimizer for better performance. We introduce CLUE (Curating Learning from Useful Experience), a framework that augments prompt optimizer for selecting which agent experiences should shape prompt optimization. We evaluate CLUE on GoBrowse, ALFWorld, AppWorld, and tau-bench. Relative to GEPA, CLUE improves held-out success by up to 7.4% on GoBrowse, 9.4% on ALFWorld, 20.7% and 18.4% on AppWorld Test-N and Test-C, respectively, and 8.6% on tau-bench.