DEALTRACE: Localizing and Intervening on Failures in Agentic Financial Reasoning
Abstract
Reviewing a private-equity deal is long, multi-step knowledge work: an analyst reads a narrative memorandum against the Excel model meant to support it, rebuilds the forecast behind the valuation, and writes a recommendation a committee acts on. LLM agents are now asked to do this, but existing benchmarks test pieces of it in isolation and score only the final answer. DealTrace pairs real memoranda with their financial models, grades each stage against the source it cites, and has a four-judge panel grade the writeup. Because every claim carries its provenance, we can repair or remove a stage and see what the conclusion rested on. It rests on one: a true forecast lifts every model by twenty points or more over the preceding rung, a true extraction or reconciliation by a few.