EditJudge-Bench: Auditing VLM Image-Edit Judges with Synthetic Ground Truth
Jens Parslov ⋅ Marco Schouten ⋅ Dim Papadopoulos
Abstract
Vision–language models (VLMs) are increasingly used as automated judges for image-editing benchmarks, yet their reliability is poorly understood. Auditing these judges is difficult because real edited images lack controlled ground truth and entangle the requested change with model artifacts, unintended side effects, and ambiguous degrees of success. In this paper, we introduce EditJudge-Bench, a diagnostic benchmark of 1,500 synthetic before/after image pairs rendered in Blender, with edits parameterized across nine categories derived from real user requests and exact ground truth from known 3D scene parameters. Because every edit is generated by a known intervention on a controlled scene, EditJudge-Bench evaluates VLM judges as binary verifiers of whether an instruction was correctly executed, isolating judge perception from editor quality. Across 18 judge configurations, we find substantial reliability gaps: (a) the widely adopted VIEScore template reaches only 67.1\% AUROC with GPT-4o; (b) fine-tuned judges score 50.9–65.0\%; spatial edits such as rotation and shape remain near chance; (c) and prompt-formatting choices like image–text ordering shift performance by up to 13\%. Despite using only synthetic data, judge rankings on EditJudge-Bench correlate strongly with two real-image human-preference benchmarks ($\rho$=0.93 and $\rho$=0.77), indicating that controlled audits surface failure modes relevant to real evaluation.
Chat is not available.
Successful Page Load