Who Should Evolve? Uncertainty-Aware Role Bottleneck Inference for Multi-Agent LLM Training
Abstract
Reinforcement learning from trajectory-level outcomes has become a standard approach for training role-partitioned LLM agent pipelines, in which multiple specialized roles such as an orchestrator, a retriever, and a synthesizer collaborate to solve complex reasoning tasks. Upon trajectory failure, existing training methods do not explicitly determine which role should receive updates, applying outcome signals to all roles indiscriminately or diffusely across the pipeline. However, this paper identifies that failures can often be traced to a primary responsible role, which we term the bottleneck role. This suggests that updating all roles may contaminate non-bottleneck roles with irrelevant updates. Therefore, an intuitive solution is to exclusively update the bottleneck role for each failed trajectory. Unfortunately, accurately identifying this role is highly challenging, as we show that roles exhibiting visible symptoms of failure are often merely downstream victims of errors originating from other roles, rather than being the true bottleneck. To tackle this challenge, we propose Role Inference for Selective Evolution (RISE), the first method that formulates per-trajectory update-target selection as role bottleneck inference. Specifically, RISE leverages Beta-binomial lower-confidence scores to estimate role responsibility, effectively discounting uncertain evidence and suppressing updates when attribution confidence is insufficient. Furthermore, RISE decouples the visible symptom locus from the support-constrained responsibility target prior to applying selective role-level training. Empirically, across HotpotQA, 2WikiMultihopQA, and MuSiQue with Qwen2.5-7B and Llama3-8B backbones, RISE consistently outperforms the strongest learning-based baseline, with up to 11.1% relative F1 and 14.1% relative EM improvements over GiGPO on MuSiQue with Llama3-8B. Code is available at: https://anonymous.4open.science/status/RISE-C0F0