Agents' Last Exam
Abstract
Recent AI systems achieve strong results on many benchmarks, yet these gains have not translated into economically meaningful deployment across professional domains. We argue that this gap is largely an evaluation problem: existing benchmarks rarely measure sustained performance on real, economically valuable workflows. We introduce Agents’ Last Exam (ALE), a benchmark for evaluating generalist computer-use AI agents on long-horizon professional tasks with verifiable outcomes. ALE is developed in collaboration with 250+ industry experts and organized around a taxonomy grounded in O*NET / SOC 2018. The benchmark spans 55 subdomains across 13 industry clusters and currently contains 1,490 task instances sourced from authentic professional workflows. Tasks require integrated GUI interaction, shell execution, software operation, and long-horizon planning inside real computing environments. To support scalable evaluation, ALE uses deterministic deliverable-based scoring and structured rubric verification rather than open-ended human judgment whenever possible. Experimental results show that current frontier agents remain far from saturation: across mainstream harness and backbone configurations, the average full-pass rate on the hardest tier is only 2.6%. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for measuring the gap between benchmark success and GDP-relevant impact.