Ledger: A Path-Validated, Database-Grounded Benchmark for Enterprise Web Agents
Xiao Yang ⋅ Mo Sha ⋅ Yiran Li ⋅ Chengjin Tian ⋅ Sheng Wang ⋅ Fangyuan Zhou ⋅ Feifei Li
Abstract
Web agents promise to automate consequential operations on enterprise web systems, yet the benchmarks used to evaluate them score correctness on surface-level proxies: DOM/URL predicates, action-trace matchers, or vision-language verdicts on a closing screen. These proxies are inadequate for enterprise resource planning (ERP) systems, where every action commits to a transactional ledger and the criterion of correctness is what the ledger now records, not what the page now renders. We argue that ERP web operations should be evaluated against what a run actually wrote to the relational backend, with the server-side operation trace recorded alongside the outcome to provide the audit provenance that enterprise governance requires. We introduce Ledger, a path-validated, database-grounded benchmark for browser agents on a production-grade open-source ERP: 424 oracle-certified task instances across 53 functional modules, scored by an Atomic Edit Set validator (field-level $F_1$ against a PostgreSQL oracle, under a typed normalization function) paired with a Canonical Operation Tracer (server-side instrumentation of the ERP’s JSON-RPC dispatch). Each episode runs under pg_restore snapshot isolation; no language-model judge is invoked; three formal guarantees (score determinism, episode isolation, tracer completeness) make scores auditable and exactly replayable. A cross-paradigm evaluation of seven contemporary browser agents shows Ledger discriminates architectures sharply, with task-success rate spanning 81.6% down to 13.7% across the roster while long-horizon enterprise workflows remain far from saturated.
Chat is not available.
Successful Page Load