Esmeralda: A Benchmark for Enterprise Service Management Agent Readiness
Abstract
Agents are increasingly being deployed to perform meaningful work in enterprise systems, yet existing benchmarks capture only a narrow slice of this work, and often rely on proxy environments that emulate real-world platforms. We introduce Esmeralda a benchmark of hard ServiceNow tasks executed directly on live platform instances and derived from expert-authored certification labs. Tasks reflect realistic enterprise workflows, with some executable through APIs, others through the UI, and some requiring both. We develop an automated multi-agent workflow for producing complete benchmark samples, in which coding agents collaboratively generate, execute, repair, and review task artifacts against fresh ServiceNow instances. Each sample includes a setup routine, an independent state-based verifier, and a golden reference trajectory, and must pass both deterministic consistency checks and LLM-based audits before acceptance. Task instructions are human-written, and we additionally conduct independent human verification in the ServiceNow UI to confirm the intended state changes. Across 212 tasks spanning 18 certification courses, all six frontier agents we evaluate struggle, the strongest solving just over half. Cost is poorly aligned with capability: the two strongest agents differ by less than a percentage point in success yet by 2.75x in price per attempt, while no agent under $0.80 per attempt solves more than a third of the benchmark.