Spend Only What You Need: Defect-Aware Residual Coverage for Efficient Multi-Agent Reasoning
Abstract
Multi-agent orchestration over large language models is increasingly used to improve reasoning accuracy, especially on problems that benefit from critique, verification, specialization, or consensus, but these gains often come with inference costs that are not explicitly controlled. Many systems spend similar compute on straightforward queries and genuinely ambiguous or expert-level ones. Existing methods typically follow fixed collaboration protocols, compile task-level workflows offline, route queries among models or collaboration modes, or use learned controllers whose state does not explicitly identify which reasoning defects remain unresolved. As a result, they do not directly decide during inference whether the current multi-agent trajectory still justifies another agent call. We introduce Defect-Aware Residual Coverage (DARC), a training-free inference-time controller for adaptive multi-agent reasoning that represents the current trajectory using four residual defect dimensions: answer uncertainty, claim-level contradiction, verification failure, and aspect under-coverage. DARC selects the next agent by maximizing cost-normalized submodular marginal coverage over the remaining residuals. Each candidate agent receives a role-induced capability profile from frozen embeddings of its natural-language role description, enabling DARC to operate without task-specific supervision, learned routing parameters, or fine-tuning. Across five reasoning and multimodal suites, MMLU-Pro, GPQA-Diamond, LiveBench, MMMU-Pro, and HLE, DARC achieves the strongest accuracy under the main shared heterogeneous six-agent pool while using substantially fewer tokens and agent calls than recent workflow, adaptive orchestration, and cost-aware routing baselines, with additional large-pool experiments showing consistent scaling behavior. It also outperforms the single-call frontier reasoning models tested, showing that residual-aware selective invocation can improve both answer quality and inference efficiency.