Westworld Finance Diligence Bench: Evaluating AI Agents on Long-Horizon Enterprise Work
Robert Alward ⋅ Alina Hyk ⋅ Victoria Knapp-Pérez ⋅ Wyatt Marshall
Abstract
Enterprises increasingly deploy language-model agents against complete business processes, while agent benchmarks continue to grade isolated questions. We present Westworld Finance Diligence Bench (WFDB), an expert-curated benchmark of 88 independent, long-horizon problems drawn from a company-acquisition deal process. Each problem is a multi-artifact workflow in its own right: runs average over 150 steps and the longest exceed a thousand, spanning a data room, email, chat, and an office suite, on anonymized documents from real private transactions. Deliverables are graded by expert-weighted deterministic checks and rubrics, and every problem goes through a six-check quality-assurance pipeline that includes an adversarial evaluation of the verifiers. Across seven frontier models in fourteen model-harness configurations, mean scores top out at $0.497 \pm 0.030$, below half the available credit. Allowing code execution helps every model, but by amounts ranging from $+0.002$ to $+0.149$, so harness sensitivity is a property of the model rather than of the task. Step-level labeling of every run shows two distinct failure axes: instruction loss over the horizon carries 42.4 percent of long-horizon capability failures, while omitted required work, wrong analytical method, and copying instead of deriving together carry 75.2 percent of domain failures.
Chat is not available.
Successful Page Load