SWO: Structured Workflow Optimization for Self-Improving AI Agent
Abstract
AI agents are becoming quite powerful, yet deploying a reliable agent still requires substantial human engineering effort. When an agent fails, the evolving cycle of an AI agent is slow due to human needs for inspecting execution traces, identifying the root cause of failure cases, deciding which part of the agent should be modified (e.g., scaffold, prompt, routing logic, or code), and retesting the agent's performance iteratively. As the AI-agent system becomes more complex, this manual eval-repair loop becomes the main bottleneck for agent deployment. Recent approaches for evolving agents mainly focus on prompt optimization and goal-driven optimization with scalar reward. However, we observe that this is insufficient for reliable optimization, especially in multi-agent systems, where failures may arise from multiple components across the entire agent workflow. We introduce \emph{Structured Workflow Optimization (SWO)}, an evaluation-driven optimization framework that enables a meta-agent to iteratively diagnose, repair, and improve an agent workflow. SWO provides the optimizer with three forms of structured context: (1) failure structure, which separate symptom taxonomy from candidate root causes analysis; (2) an explicit editable-surface representation, which exposes the heterogeneous components of an agent (e.g., prompts, tools, workflow scaffolding, and executable code); and (3) cross-iteration optimization memory, which records previous hypotheses, modifications, evaluation outcomes, and rejected fixes to prevent redundant or regressive updates. At each iteration, the framework evaluates the current agent, analyzes its failure distribution, proposes a prioritized intervention, modifies the corresponding workflow component, and validates the change against the evaluation suite. By leveraging these three structured contexts, SWO can localize failure sources efficiently, select the editable surface from wider space, and avoid redundant or regressive updates across iterations. We evaluate SWO on public interactive-agent benchmarks. Across the evaluated settings, our method consistently improves task success over the initial agents and substantially outperforms unconstrained native-agent self-optimization, while requiring fewer optimization tokens. Our results suggest that reliable agent self-improvement depends not only on stronger optimization models, but also on explicitly structuring the evidence, intervention space, and optimization history available to the optimizing agent.