WordEval: Evaluating Word-Native Operation Fidelity in Document Editing
Abstract
Microsoft Word is a central environment for office automation, yet evaluating whether an agent has correctly edited a Word document remains difficult. Unlike plain text generation, Word editing changes a structured document object, where formatting, lists, fields, layout, and object properties may all be part of the requested effect. Existing evaluation protocols based on whole-document text similarity, rendered visual similarity, or LLM-as-a-judge scores cannot reliably determine whether the requested Word-native effect is realized as editable document state or whether relevant surrounding state is preserved. We introduce WordEval, a benchmark for evaluating Word-native document editing from natural-language instructions. Each case consists of an input .docx file and an editing instruction; the evaluator compares the agent’s output with a reference document using deterministic checks over Word-readable object properties. WordEval is constructed by design from planned editing units, rather than sampled from arbitrary documents: each unit specifies an input precondition, target object, operation, parameters, reference-generation procedure, and oracle specification. The current release contains 642 cases, including 546 atomic cases over 231 Word commands and 96 compositional cases instantiated from 27 composition templates. Its analysis layer covers 140 unique object-property pairs across 65 Word COM objects and 15 document subsystems. The compositional cases are linked to corresponding atomic cases for their component operations, allowing analyses of whether failures arise from individual operations or from combining multiple operations in one document. By combining controlled document construction, reference-based evaluation, and command-grounded deterministic oracles, WordEval provides a fine-grained testbed for measuring Word-native document automation.