Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
Abstract
Multi-agent LLM systems improve reasoning and tool use via role specialization, yet reinforcement learning (RL) post-training for such systems remains unstable and underexplored. We theoretically pinpoint a key source of instability when extending group-based RL to cooperative multi-agent LLM systems: under GRPO-style optimization, a global normalization baseline can mismatch heterogeneous agents' reward distributions, inducing gradient-norm instability. Based on this finding, we propose Dr. MAS, a simple and stable RL recipe for cooperative multi-agent LLM systems. Algorithmically, Dr. MAS normalizes advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes the training. Systemically, Dr. MAS provides an end-to-end multi-agent RL framework with scalable orchestration, flexible per-agent LLM serving and optimization, and shared resource scheduling of actor backends. Unlike single-actor frameworks such as veRL, Dr. MAS natively supports multiple heterogeneous LLMs over a unified GPU pool. Lifecycle-aware backend management and dynamic dispatch release inactive models' GPU memory and schedule active agents on demand, enabling hardware-efficient co-training. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search with Qwen2.5 and Qwen3. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6\% avg@16 and +4.6\% pass@16 on math, and +15.2\% avg@16 and +13.1\% pass@16 on search) while largely eliminating gradient spikes. It also remains highly effective under heterogeneous agent-model assignments while improving efficiency. Code is available at https://anonymous.4open.science/r/DrMAS-0787.