RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with LLMs
Abstract
Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning capabilities with external environments. However, sparse terminal rewards hinder fine-grained, state-level optimization. Although process reward modeling offers a promising alternative, training dedicated reward models often entails substantial computational costs, risks reward hacking, and faces annotation bottlenecks. To address these challenges, we introduce RewardFlow, a lightweight method for estimating state-level rewards for agentic reasoning. RewardFlow leverages the intrinsic topological structure of states within trajectories by constructing state graphs. This enables topology-aware graph propagation to estimate state-wise contributions to success, yielding principled, annotation-free state-level rewards. As dense rewards for RL optimization, RewardFlow substantially outperforms prior RL baselines across four agentic benchmarks, with average success-rate gains of +6.2\% on text-based benchmarks and +29.7\% on visual reasoning over the strongest baseline across three model scales, and over +10\% accuracy improvement on DeepResearch, while demonstrating superior robustness and training efficiency.