Adaptive Auto-Harness: Sustained Self-Improvement on Open-Ended Task Streams
Abstract
Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, and supporting infrastructure from execution feedback, but are typically evaluated on fixed offline benchmarks. Real deployments instead present open-ended task streams: histories grow without a fixed endpoint, heterogeneous tasks require different harnesses, and distributions shift. Under that pressure a single densely updated harness turns brittle, with accuracy peaking early and then declining, and three published systems end below the solver they started from. We cast harness construction as regret minimization against an oracle harness and decompose the gap into an evolution loss and an adaptation loss. We then introduce Adaptive Auto-Harness, whose stateful multi-agent evolver targets the evolution loss, with two further bounded axes, solve-time routing over a harness tree and a human channel for signals history cannot contain. Across prediction-market, security-competition, and event-forecasting streams it outperforms five auto-harness baselines. A mechanism analysis shows that the ranking on every stream is decided by two evolver capabilities rather than by evolver budget: global-signal mining, reading the whole history for a cost no single trajectory reveals, and consolidation, folding edits into one coherent artifact instead of accumulating case-scoped fixes. Sustained self-improvement therefore comes from having both capabilities, not from a larger evolver budget or from tuning the harness to each benchmark.