GeoGenBench: Symbolically Verified Generation of Geometry Diagrams
Abstract
Typically, geometry benchmarks ask models to read and interpret diagrams. In GeoGenBench, we ask the models to draw. Each of the 801 prompts describes a plane-geometry construction and comes with the checks a correct drawing has to pass. The model writes TikZ. We compile the code, read the coordinates back out, and run every check exactly, so no person and no vision model has to judge the drawing. We test six frontier models. Asking for a typed intermediate representation instead of raw TikZ helps every one of them, by 2 to 25 points, and helps the weakest the most. With that scaffold, the three best models land within a point of each other at prices that differ sevenfold. The grader has one known gap. Every check it runs is exact, but the checks do not always cover everything a prompt asks for. When two people graded 200 outputs by hand, the grader never failed a correct drawing, but it passed about one drawing in six that a person would not. We release the dataset, the templates, the grader, and the leaderboard. The remaining work is making each template's checks cover every requirement in its prompt.