OpsBench: A Reliability-Oriented Benchmark for Open-Weight Agents on Internal Enterprise Operations
Willy Fitra Hendria
Abstract
The $\tau$-bench family evaluates enterprise agents only on customer-facing verticals (retail, airline, telecom, banking knowledge retrieval) using a dialogue-driven, tool-agent-user methodology with reliability-oriented pass$^k$ scoring. Separately, recent internal-operations benchmarks such as EnterpriseOps-Gym, Top of the CLASS, and EnterpriseBench cover IT, HR, and finance/administrative tasks, but none evaluate against a simulated-user dialogue loop, and results are reported primarily on paid frontier models. We introduce OpsBench, which combines $\tau$-bench's dialogue-driven pass$^k$ methodology with an internal IT-helpdesk domain (identity verification, hardware-approval thresholds, escalation policy, state-inferred fabrication resistance) spanning 18 scenarios across six policy dimensions, evaluated entirely on eight open-weight models run on local consumer hardware at zero monetary cost, a reproducibility profile distinct from prior internal-ops benchmarks. We find that (1) reliability under repetition remains the binding constraint even at this modest scale: only five of eight models show any pass$^3$ reliability at all, and the best model reaches only 50% pass$^3$ despite an 80% single-run pass rate; (2) token cost is a poor predictor of accuracy at matched token budget (two models with near-identical total token usage differ by 46 points of pass rate and 39 points of pass$^3$), even though total parameter count itself correlates positively with accuracy in our sample (Spearman $\rho=0.63$, not significant at $n=8$), so the cost-accuracy mismatch is about architecture and training, not a claim that scale doesn't matter; and (3) a model nominally trained for tool use can still fail purely from a serving-stack format mismatch, a failure mode distinct from, and easily conflated with, a lack of capability. We release the full suite (mock environment, policy specifications, and evaluation harness) to support further work on internal-operations agent evaluation, at an anonymized mirror: https://anonymous.4open.science/r/opsbench-C68D/.
Chat is not available.
Successful Page Load