Beyond Outcome Rewards: Process-Aware Optimization for Search Agents
Abstract
Search agents based on large language models address complex information-seeking tasks by iteratively issuing queries, retrieving evidence, and integrating information. Effective training of such agents requires supervision beyond final-answer correctness, as outcome-level rewards cannot distinguish efficient evidence acquisition from redundant or misdirected search. Recent process-level methods introduce intermediate supervision, but typically apply uniform signals across trajectories. However, such uniformity overlooks the state-dependent nature of process quality, where the value of an intermediate decision depends on the current evidence state. This motivates a systematic analysis of how process quality varies across trajectories with different evidence states and final outcomes. To this end, we analyze search trajectories on BrowseComp-Plus and xBench, revealing two consistent patterns: successful trajectories often suffer from post-alignment inefficiency (redundant searches), whereas failed trajectories typically struggle with misdirected exploration and insufficient coverage. Motivated by these findings, we propose Process-Aware Grouped rEward optimization (PAGE), a framework for search agents that provides a concrete instantiation of state-dependent process supervision. PAGE partitions trajectories according to final outcomes, assigns differentiated process rewards to successful and failed groups, and integrates them with outcome supervision at either the reward or advantage level. Experiments on three backbones across six multi-hop QA and web-search benchmarks show that PAGE consistently improves answer accuracy over outcome-only baselines.