Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Zhou Yu ⋅ Bin Bi ⋅ Shiva Kumar Pentyala ⋅ Shubham Mehrotra ⋅ Sougata Chaudhuri ⋅ Shilpa Bhagavath ⋅ Zeyuan Chen ⋅ Ran Xu ⋅ James Zhu ⋅ Sitaram Asur ⋅ Phil Mui
Abstract
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success, and automated harness evolution has recently proven highly effective at enabling smaller models to perform well on domain-specific tasks at a fraction of the cost of frontier models. Since both the harness and the model's weights shape an agent's behavior, and fine-tuning methods such as LoRA are a widely adopted practical model lever for small and mid-sized models, we ask how these two levers should be combined. Using seven enterprise agentic benchmarks, we first evolve a harness with the weaker model and realize its gains; we then find that a stronger expert model often makes even better use of the evolved harness, suggesting that the weaker model could learn from the expert to close the remaining performance gap. However, the natural next step of teaching the weaker model using the expert's trajectories under the evolved harness surprisingly backfires: imitating the expert causes the weaker model to regress on all seven tasks ($-4$ to $-30$ points), a result reproduced across two model families (Qwen3-Coder and Gemma~4)\footnotemark[1], even though the identical procedure helps under the unevolved baseline harness. Our analysis reveals that, under imitation, the weaker model \emph{does} acquire the expert's knowledge and makes greater use of the harness scaffold, but its fit to the harness is impaired. The weaker model adopts the expert's planning strategy without the competence to execute it and no longer fits the harness that was evolved around its native planning style. To effectively introduce a teaching signal, we instead develop an on-policy expert-correction pipeline, automated end-to-end by a meta-level MLE agent: it localizes the failing turn in each of the weaker model's own rollouts and has the expert rewrite only that turn, preserving the model's planning style. This approach learns from the expert without breaking harness fit, thereby combining the benefits of both harness evolution and model adaptation. We identify and resolve a source of contention between harness evolution and model-weight adaptation, yielding a recipe can be incorporated into a co-evolution loop. Our results shed light on how to jointly and economically co-evolve harnesses and model weights to achieve strong performance on domain-specific enterprise tasks.
Chat is not available.
Successful Page Load