Evaluating LLM Agents on Multi-Modal Professional Finance Workflows
Abstract
While frontier models now match experts on some closed-form reasoning problems, professional finance work remains open-ended, iterative, and spread across many documents. Most finance or multi-modal benchmarks still measure only isolated units of a real workflow. We introduce FinXVal, a benchmark of 37 tasks across five task families that test agentic models along three end-to-end task dimensions. Our tasks have cross-document, multi-modal inputs, with ambiguous authority chains that require the agent to interpret and revise native Office artifacts. Tasks are built and independently reviewed by finance domain experts and graded against 1,677 atomic rubric criteria. Across twelve frontier models and four agent harnesses, no model averages above 49 out of 100. FinXVal shows that judgment-intensive, file-native finance work remains an open problem for agentic systems.