A Diagnostic Benchmark for Layered Layout and Template-Variant Reasoning
Jaejung Seol ⋅ Haonan Zhu ⋅ Elad Hirsch ⋅ Adrienne Deganutti ⋅ Purvanshi Mehta
Abstract
Graphic-design layouts are structured, multi-layered artifacts: a single composition specifies the position, type, z-order, and styling of dozens of components, and a single template typically begets a family of sibling variants that share its structural theme but differ in content, palette, or imagery -- properties that flat-canvas layout benchmarks discard. We introduce \textsc{GraphicDesignBench} (GDB), a diagnostic benchmark of $16$ tasks along two axes: \emph{spatial composition} ($8$ understanding + $4$ generation) and \emph{template variants} ($3$ understanding + $2$ generation), grounded in a new dataset of $1{,}148$ real-world templates with full component hierarchy, per-element styling, and sibling groupings. Across four frontier VLMs (GPT-5.4, Gemini-3.1-Pro / Flash-Lite, Claude-Opus-4.6) and two image generators (GPT-Image-1.5, Gemini-3.1-Flash-Image), GDB exposes concrete and dissociated gaps: component detection sits an order of magnitude below natural-image baselines, layer-order reasoning decouples from other spatial skills, partial completion collapses from single- to multi-element settings, and a simple font-Jaccard baseline matches or exceeds frontier VLMs on sibling matching and clustering. We release the dataset, tasks, metrics, and per-model outputs to track where current models are, and are not, ready to act as design collaborators.
Chat is not available.
Successful Page Load