MULTILOOP: : Evolving Multi-Agent Harnesses for Complex Information-Seeking Tasks
Abstract
The performance of an LLM-based multi-agent system depends not only on the underlying model but also on the harness, including prompts, tools, and sub-agent orchestration logic. As agentic systems grow in complexity and produce increasingly long trajectories, manually inspecting failures to improve the harness becomes prohibitive, and even more cumbersome in multi-agent systems that include multiple sub-agent trajectories. We introduce MULTILOOP, an automated research loop in which coding agents iteratively analyze multi-agent trajectories, identify common issues, implement harness modifications, and evaluate downstream performance to guide subsequent optimization rounds. Through a case study on a multi-agent harness for complex information-seeking tasks, we show that MULTILOOP produces improvements addressing issues such as hallucinations, query formulation errors, inadequate error handling and missing tools, among others, resulting in a 4.6% improvement over the baseline harness on the challenging DeepSearchQA benchmark with only 10 training samples, paving a path towards autonomous multi-agent harness optimization.