OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
Abstract
Building capable visual web agents demands precise grounding, long-horizon reasoning, and robust interaction with dynamic, real-world websites. Despite rapid progress, the strongest systems remain largely proprietary, while open agents still depend heavily on supervised post-training over large collections of curated web trajectories. This dependence creates a major scalability bottleneck: high-quality demonstrations are expensive to collect, and static datasets offer limited coverage of the diverse, ever-changing open web. Although online reinforcement learning (RL) has shown promise for text-based agents, its potential for training visual web agents directly on live websites remains largely unexplored. In this paper, we introduce OpenWebRL, an open framework for training visual web agents with online multi-turn RL on real websites. OpenWebRL covers the full training pipeline, including task selection, supervised warm-starting, live-browser execution, multimodal context management, trajectory success judging, and efficient multi-turn policy optimization. Using this framework, we train OpenWebRL-4B, which establishes a new open-source state of the art on challenging live-web benchmarks. With only 0.4K initialization trajectories and 2.2K open-ended training tasks, OpenWebRL-4B achieves 67.0% success on Online-Mind2Web and 64.0% on DeepShop, outperforming prior open agents of similar or larger scale and remaining competitive with proprietary systems including OpenAI and Gemini CUA. Beyond strong benchmark performance, OpenWebRL systematically identifies the key design choices that make online RL effective for visual web agents. Overall, our work offers a practical path toward building more capable, reproducible, and cost-efficient open web agents. We will release our training data, models, and code to support future research.