Learning When to Collaborate: Selective Multi-Agent Medical Reasoning via Uncertainty-Aware Routing
Jiacheng Hou ⋅ Xueliang Cui ⋅ Dan Lu ⋅ Ruxin Wang
Abstract
Multi-agent systems have emerged as a promising paradigm for multimodal medical reasoning, but existing methods rely on two assumptions that hinder practical deployment: (i) always-active dense collaboration among multiple agents, and (ii) dependence on large cloud-based models that raise privacy and efficiency concerns. In this paper, we challenge both assumptions and propose SMART-Med, a fully local multi-agent method that treats collaboration as a learnable decision rather than a default behavior. Our key insight is that not every medical query requires costly multi-agent deliberation: many can be reliably solved by a single well-trained agent, while only uncertain cases benefit from expert collaboration. Based on this insight, SMART-Med introduces an uncertainty-driven routing mechanism that adaptively decides when to answer directly and when to recruit a small subset of expert agents, effectively converting dense multi-agent reasoning into selective, query-adaptive coordination. To make this routing reliable, we curate a new multimodal medical dataset MedVR-3K with high-quality reasoning traces and a difficulty hierarchy, and then fine-tune local agents using GRPO with two complementary rewards: a certainty-aware accuracy reward and a vision-language reward model for correctness and reasoning quality respectively. Experiments on four medical visual question answering benchmarks show that SMART-Med achieves competitive performance compared to existing MAS approaches while substantially reducing inference cost (reducing inference time by more than 3.8$\times$ and token usage by 72\%). Moreover, after fine-tuning on MedVR-3K, SMART-Med yields an average improvement of 6.7\% across Slake, PathVQA, VQA-RAD, and PMC-VQA. Our results suggest that selective collaboration can be an efficient yet effective alternative to dense collaboration and enable practical, privacy-preserving deployment of multi-agent medical AI.
Chat is not available.
Successful Page Load