DAGent: Evaluate-then-Grow Planning for Deep Research Agents
Abstract
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as intermediate findings emerge. Directed acyclic graph (DAG)-based multi-agent systems are well suited to this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents typically instantiate a task-level plan before execution and then repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions often waste computation on branches that should not have been planned in the first place. To bridge this gap, we propose DAGent, a DAG-based multi-agent framework that introduces Evaluate-then-Grow incremental planning. Instead of committing to a full DAG upfront, an Orchestrator grows the task graph one batch at a time, conditioning each new expansion on confidence and uncertainty signals from completed nodes. To support long-horizon evidence use without overloading each sub-task, DAGent maintains a hierarchical context layer that propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The cleanly recorded DAG topology in turn admits structural RL signals that an outcome-only recipe cannot define; we instantiate this with DAGRPO, a GRPO adaptation that injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, with the lead replicating across four open-source backbones from four vendors and extending to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO further improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison further shows that incremental, evidence-conditioned planning recovers accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart, supporting the view that DAGent's gains come from more targeted evidence expansion rather than additional computation. Anonymized code is available at \url{https://anonymous.4open.science/r/DAGent-F2D6}.