RegFPO: Leveraging Regret-Based Feedback for Agentic Post-Training with Tree-based Sampling
Abstract
We introduce REGFPO, a policy-optimization method that uses decision-level regret-based feedback for credit assignment. We derive decision-level training signals by decomposing regret into action suboptimality gaps. REGFPO estimates these gaps from sampled rollout subtrees rooted at selected decision points. We theoretically show that (Tree-)GRPO return-based updates can decrease the probability of an optimal action when its action value under the current policy is lower than that of a suboptimal action, whereas REGFPO constructs a decision-level advantage by comparing each action’s estimated value with the highest estimate at that decision. Experiments on finite tree-structured MDPs show that these different decision-level advantages change which actions remain likely to be sampled during training. On four agentic benchmarks spanning robotic manipulation, web interaction, browser control, and embodied reasoning, the evaluated REGFPO configurations attain the highest scores among the compared methods, with margins of 2.1–29.1 points over the best non-regret result across the reported evaluation splits.