SAPO: Step-Level Skill-Augmented Policy Optimization for Multi-Turn LLM Agents
Abstract
Large Language Model (LLM) agents have shown strong results on multi-turn tool-use tasks, yet they typically operate in isolation during training, failing to leverage skills accumulated across episodes. Existing experience-augmented methods address this by organizing past trajectories into retrievable libraries, but they retrieve skills only once based on the initial task description and hold them constant throughout the episode. In multi-turn settings where observations change at every step, this static retrieval becomes increasingly mismatched as episodes progress. We propose SAPO (Step-Level Skill-Augmented Policy Optimization), a reinforcement-learning framework that retrieves relevant skills at each decision step conditioned on the current observation. SAPO operates through three components: (i) step-level observation clustering that groups structurally equivalent environmental states for efficient cluster-indexed retrieval; (ii) a self-evolving skill bank that distills successful strategies and failure patterns through score-based admission and rate-limited extraction; and (iii) policy optimization with step-level credit assignment for fine-grained advantage estimation across multi-turn episodes. The skill bank evolves alongside the policy through semantic analysis rather than gradient updates. On long-horizon multi-turn agent benchmarks (ALFWorld, WebShop, and seven search-augmented QA tasks), SAPO achieves 93.5\% on ALFWorld, 76.3\% success on WebShop, and 60.9\% average across QA tasks, outperforming both standard RL and prior skill- and experience-augmented baselines. Our code is available at \url{https://anonymous.4open.science/r/slea-rl-4D1E/}.