Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
Lunjun Zhang ⋅ Ryan Chen ⋅ Bradly Stadie
Abstract
Building agentic systems that can autonomously self-improve from experience is a longstanding goal of AI. Large language models (LLMs) today primarily self-improve via two mechanisms: self-reflection for context updates, and reinforcement learning (RL) for weight updates. In this work, we propose Evolutionary System Prompt Learning (E-SPL), a method for jointly improving model contexts and model weights. In each RL iteration, E-SPL samples trajectories under multiple system prompts in parallel, then jointly applies RL updates to weights and evolutionary updates to system prompts via LLM self-reflection. E-SPL encourages a natural division between declarative knowledge encoded in prompts and procedural knowledge encoded in weights. Most notably, in an easy-to-hard generalization setting (AIME $\rightarrow$ BeyondAIME), E-SPL improves RL success rate from 38.8% $\rightarrow$ 45.1%. E-SPL also improves RL on AIME 2025 (56.3% $\rightarrow$ 60.6%), HMMT 2025 (50.0% $\rightarrow$ 52.7%), and agentic search (44.2% $\rightarrow$ 48.6% on gpt-oss-120b). Across all settings, E-SPL outperforms both RL-only and evolution-only baselines, demonstrating that weight updates and context updates are deeply synergistic and can together yield gains in generalization that neither achieves alone.
Chat is not available.
Successful Page Load