Challenge-Aware Knowledge Recruitment with Meta-Agent Supervision for Medical Multi-Agent Reasoning
Abstract
Large language models have shown strong performance on standard medical question-answering benchmarks, yet deteriorate significantly on more complex medical reasoning tasks. Multi-agent debate frameworks, in which a team of specialist agents discuss before reaching a consensus answer, represent a promising approach, but existing systems often fail to outperform strong single-model prompting baselines on such tasks. We argue that current systems share a common limitation, as they select medical experts and retrieve external knowledge based solely on the question text, without accounting for the expertise and knowledge gaps that only emerge as the discussion progresses. We propose a framework of challenge-aware adaptive recruitment, in which an external CHALLENGER agent evaluates the reasoning trace between discussion rounds, identifying specific gaps in logic or factual knowledge. This signal drives two adaptive mechanisms. Knowledge recruitment triggers targeted retrieval after a factual gap has been concretely exposed. Expert recruitment introduces domain-relevant specialists mid-discussion when an expertise gap is identified. This design separates the question of when to recruit from whether resources are available. We evaluate the framework on the MEDAGENTSBENCH hard subset (862 questions across nine biomedical datasets). For knowledge recruitment, standard upfront retrieval augmentation was found largely ineffective on hard medical questions, while challenge-aware in-between retrieval improved performance on 8 of 9 datasets, adding 6 percentage points over the Challenger-only baseline. Both the timing of retrieval and the specificity of its trigger contribute independently, and a live web search experiment further confirms this timing advantage is architectural rather than corpus-dependent. For expert recruitment, we first show that self-triggered adaptation is unreliable and tends to be harmful. By contrast, Challenger-triggered expert recruitment improves performance consistently, raising MedQA accuracy under GPT-4.1-mini from 46\% to 54\%, with positive direction across all generalization datasets tested. We further show that naive combination of both channels degrades performance. Both proposed mechanisms individually achieved strong results on the benchmark, establishing that the question of when and why external resources are recruited is a decisive factor in multi-agent medical reasoning.