Learning to Surpass: Training Tool-Using Agents with Anchored Feedback
Abstract
Training tool-using agents for planning is challenging when high-quality demonstrations exist only as final outcomes without the intermediate tool traces that produced them. A common response is to impute traces and apply imitation learning, but this approach is indirect and structurally limited to matching the reference rather than improving upon it. We propose anchored comparative feedback: rather than imitating a reference plan, we train agents to surpass it under a comparative LLM judge while satisfying hard constraints. The anchor provides a stable optimization target, and comparing two identically rendered plans neutralizes surface-level judge biases. On TripTailor, a tool-augmented travel planning benchmark, anchored GRPO achieves 80.1% final success rate compared to 43.7% for imitation learning on imputed traces, with gains on both constraint satisfaction and LLM-judged preference that persist across judge families, rubrics, and a held-out reward model. Gains generalize beyond TripTailor, with a fivefold improvement on the held-out TravelPlanner benchmark. These results suggest that anchored comparative feedback offers an effective approach for learning transferable tool-using planners from outcome-only supervision.