From Static Filters to Adaptive Gates: Co-Evolving Synthetic Validation Gates and Harnesses in Meta-Harness Optimization
Abstract
Agent harnesses, the code that defines what information gets stored, retrieved, and presented to the model, has become a target of enhancement. Harness engineering unlocks a new dimension of improving Large Language Model (LLM) performance beyond pre-training and post-training. Recent work has shown that autonomous agents are capable of evolving harnesses and agent skills to improve agent performance across various tasks. SkillOpt and SkillOpt-Lite shown that performing skill evolutions with validation gates to filter and select candidate skills improves evolution performance. However, these frameworks derive validation tasks from the existing task set, an approach that presents two primary drawbacks. First, it diminishes the pool of tasks available for evolution, which exacerbates performance issues in sample-limited scenarios. Second, using static tasks from the taskset as validation gates provides limited signal for filtering candidates prior to evaluation; we argue that this gating stage should instead play a more active, dynamic role in driving candidate evolution. To address these limitations, we propose an improvement to Meta-Harness, adding a co-evolving synthetic task validation gate tailored to probe harness adherence to specific behavioral constraints. By distilling these constraints from past trajectories to evolve targeted gates, the validation stage actively steers candidate evolution toward optimal policies. On coding tasks, evaluation on TerminalBench-2 shows that our framework improves average task score performance of best candidate by 3.8\% over Meta-Harness. Our findings also highlight the importance of co-evolving synthetic tasks together with candidate harnesses to achieve the best performance.