Role-Decomposed Multi-Agent LLM Training with Cross-Agent Learning Signals
Abstract
Agentic search systems must coordinate evidence acquisition and response generation, yet existing approaches either couple both roles under a single agent objective or decompose them without disentangling their respective contributions to the final outcome. We introduce DAC (Divide and Cooperate), a role-decomposed training framework that, given task-specific external verification signals, trains a searcher and a generator with role-specific cross-verification rewards. DAC combines externally evaluated search sufficiency with the generator's evolving abstention decision, producing an adaptive learning signal for the searcher while avoiding generator penalties caused by insufficient retrieval. To prevent degenerate over-abstention, we further introduce hard-positive evidence augmentation, which encourages answering from sufficient but imperfect evidence. Across seven general and multi-hop QA benchmarks and two model backbones, DAC consistently outperforms strong single-agent and multi-agent baselines. Controlled evaluations show that these gains arise from improvements in both search quality and answer generation. Our results highlight the importance of explicitly assigning credit across interacting roles when training agentic search systems.