SpreadEvo: A Benchmark for Multi-Turn Evolving Intent in Artifact-Editing Spreadsheet Agents
Abstract
Spreadsheet agents are usually evaluated from a single instruction and a final workbook. Real interactions are less static: users reveal requirements over time, revise assumptions, refer back to earlier results, change representations, and sometimes abandon one path for another. Existing evolving-intent formalisms capture some of this movement as changes to a slot frame, but that view does not represent the workbook being edited or the history needed to interpret later requests. We introduce SpreadEvo, a benchmark that converts 321 workflow-level tasks from SpreadsheetBench 2 into three- to six-turn conversations spanning nine forms of slot-based, artifact-grounded, and history-dependent intent change. A structured anchor keeps dialogue-safe task information separate from verifier-only gold metadata, while turn-local grading follows the workbook as it evolves and retains the source benchmark's native evaluator at the final state. This makes it possible to measure not only whether an agent eventually finishes, but whether it preserves, revises, and recovers the user's intent along the way. In a matched 66-conversation evaluation, GPT-5.6 and Claude Opus 4.8 complete only 28.8% and 30.3% of trajectories. Verification and audit are the weakest shared interaction type, and recovery after an earlier failure occurs in only 1 of 11 eligible cases.