VisEditBench: A Benchmark for Vector-Format Diagram Editing with Visual Instructions
Abstract
We introduce a new task and benchmark dataset, VisEditBench, for understanding visual instructions in vector-format diagram editing. While recent advances have made progress in generating vector graphics from text or images, editing such diagrams remains relatively underexplored. Moreover, existing approaches predominantly rely on textual instructions, which are often unintuitive for specifying edits in structured visual content, particularly when users need to identify specific elements or indicate their spatial positions in a diagram. We propose visual instructions as a more intuitive interface for diagram editing, and construct a dataset of 694 samples with human-curated annotations overlaid on TikZ diagrams. The dataset covers multiple edit types as well as instructions with multiple editing operations. We evaluate 32 recent state-of-the-art multimodal models and find consistent performance degradation, particularly for edits involving global layout or diagram structure, compared to local operations such as element addition or deletion. Furthermore, instructions that include multiple editing operations tend to result in lower performance than single-edit cases. Even recent models struggle to achieve consistently accurate edits, highlighting the inherent difficulty of the task. VisEditBench provides a valuable benchmark for advancing research on visually grounded diagram editing and structured visual understanding.