Frontier Work: Benchmarking AI Agents on End-to-End Professional Workflows
Abstract
Enterprise agents are often evaluated on short answers or atomic software actions, but professional work ends in artifacts: reconciled workbooks, returns, clinical notes, and decision memos. We introduce Frontier Work, a frontier-level 250-task benchmark for end-to-end workflows across finance (80 tasks), tax (68), healthcare (56), and accounting (46). Each task couples a brief request and evidence workspace to required professional artifacts and a granular expert rubric. Expert-authored metadata estimate 67.2% of tasks at ten or more hours of professional work; rubrics contain approximately 53 criteria on average and as many as 199. We evaluate 19 model–system configurations with three runs per configuration–task pair, totaling 14,250 task runs. The strongest configuration, Claude Opus 5 adaptive-max, passes every rubric criterion in only 74 of 750 runs (9.87%); even at a 95% criterion threshold, only 159 runs pass (21.20%). Every configuration remains below 10% all-criteria pass@1, and the best tax result is only 4.41%. All benchmark data and experimental code will be released publicly on GitHub.