SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Abstract
Adapting language agents often means teaching them domain procedures: where to look, which tools to call, how to verify intermediate results, and how to format outputs for a grader. Existing context-level adaptation usually either hand-writes such procedures, generates them once, or lets skill artifacts grow through loosely controlled self-revision. We ask whether a skill can instead be trained as the compact external state of a frozen agent. SkillOpt is a harness-agnostic text-space optimizer for a single natural-language skill document. It runs the frozen target model on scored rollout batches, uses a separate optimizer model to reflect over success and failure minibatches, proposes structured add/delete/replace edits, ranks them under a textual learning-rate budget, and accepts a candidate only if it improves held-out selection performance. Rejected-update memory, slow cross-epoch guidance, and optimizer-side meta skill make the editing loop behave like bounded training while leaving deployment as a single best_skill.md. Across six direct-chat benchmarks covering QA, spreadsheets, documents, math, and embodied decision making, \ourmethod{} improves GPT--5.5 over no skill by 21.5 points on average, with positive gains on every task and the best measured result on five of six. Codex- and Claude-Code-style harness runs show that the same artifact remains useful in tool-backed execution. Ablations and transfer studies identify validation-gated, bounded edits and update memory as key to stable, portable skill learning.