A Subgoal-driven RL Framework for Improving Long-Horizon Web Agents
Taiyi Wang ⋅ Sian Gooding ⋅ Florian Hartmann ⋅ Oriana Riva ⋅ Edward Grefenstette
Abstract
Large language model (LLM)-based agents have emerged as powerful autonomous controllers for digital environments, spanning mobile interfaces, operating systems, and web browsers. Web navigation, for example, demands handling dynamic content and long action sequences, making it a particularly complex task. Existing LLM-backed agents exhibit weakened long-horizon planning abilities during RL fine-tuning, where sparse and delayed rewards from terminal Outcome Reward Models (ORMs) make it difficult for agents to identify the actions that lead to success, preventing them from sustaining coherent reasoning over extended tasks. We address this with two contributions: (1) a validated subgoal generation procedure that produces reliable, monotonically calibrated progress signals between the initial state and the terminal ORM reward; and (2) MiRA ($\underline{Mi}$lestoning your $\underline{R}$einforcement Learning Enhanced $\underline{A}$gent), an RL training framework using dense, milestone-based reward signals via a learned potential critic. Starting from the open Gemma3-12B base ($6.4$%), MiRA enables the agent to reach $43.0$% on the WebArena-Lite benchmark. This performance surpasses proprietary systems such as GPT-4-Turbo ($17.6$%) and GPT-4o ($13.9$%), as well as the previous open-model state of the art, WebRL ($38.8$%). Our findings demonstrate that milestone-based reward shaping significantly boosts an agent's long-horizon abilities, paving the way for more robust, general-purpose autonomous systems.
Chat is not available.
Successful Page Load