Who’s Asking? A Role-Conditioned Benchmark for Operational Reasoning in Enterprise Agents
Abstract
Enterprise agent benchmarks test whether a model can execute the work, opening an incident, routing it, closing it, but delegating judgement raises a harder question: can a model read the state of an operation and decide what to do about it, the way the person in a specific seat would? We show this question has no role-neutral answer. The same incident and audit data, read by an Incident Manager, an Operations Manager, or a Process Owner, yields different correct answers, because the three roles differ in objective, unit of analysis, and authority to act. We present a role-conditioned benchmark on an enterprise environment where every task is solvable by a scripted oracle through the agent's own tools. Ground truth is never stored: each answer is recomputed by an executable function against a frozen, fingerprint-verified copy of the instance, so the benchmark ports to any instance without re-annotation, and a typed answer contract makes wrong-role reasoning mechanically detectable. Across four models on 14 gate-admitted instances, we found that model rankings vary by role, and a pooled-score evaluation would have certified the wrong model for real operational use. These findings highlight the need for modern operational benchmarks to evaluate agents across multiple roles, as the correct answer in enterprise service operations depends on who is asking.