Diagnosing Rule Recovery in Browser Agents: A Controlled Suite for Evaluating Partial Progress
Abstract
Browser agents are typically evaluated by whether they eventually complete a task, obscuring distinct failures in interactive rule recovery. We introduce a controlled browser-based diagnostic suite that independently varies rule identifiability, partial-progress observability, and hypothesis availability while instrumenting each submitted state. Across four agents, aggregate success concealed sharply different recovery dynamics. Under opaque feedback, agents frequently reached correct partial repairs but diverged in whether they preserved them when the task continued to fail. Trace analysis showed that some agents interpreted the unchanged error as evidence against a correct repair and reversed it, whereas others retained the repair and searched for an additional constraint. Informative feedback rescued this transition for some agents but not others, and a matched-history intervention showed that making the same partial improvement observable could shift behavior from reversion toward completion. Conversely, feedback could not reinforce a counter-prior repair that agents never attempted. These results motivate evaluations that distinguish rule inference, repair preservation, and downstream hypothesis generation.