Mission Impossible: Diagnosing and Fixing Non-Operative Instruction Following in Image Editing
Guoyizhe Wei ⋅ Feng Wang ⋅ Alan Yuille ⋅ Rama Chellappa
Abstract
Instruction-guided image editing models are mostly trained and evaluated on executable requests, yet many real user instructions are non-operative for the given image. Reliable editing therefore requires reasoning not only about how to edit, but also whether to edit and what to ignore. We introduce \bench{}, a large-scale benchmark for non-operative image editing, with 1.1M instructions over 31K images across seven categories spanning symbolic, grounded, physical, and logical non-executability, plus a 35K-sample partial-edit test set evaluating selective execution when one clause is feasible and another is not. A 700-sample internal audit confirms a 94.9\% true no-op rate with $\kappa=0.87$. Evaluating 13 systems reveals a clear action--abstention dissociation: frontier editors achieve strong standard editing quality but still over-edit more than 40\% of no-op samples. Three complementary perceptual, feature-level, and vision-language metrics further surface distinct failure modes across architectures. As a reference baseline, JUDGE-THEN-EDIT BAGEL—a feasibility-aware BAGEL variant trained on a 150K mixture of full-edit, partial-edit, and no-op data—reduces over-edit rate from 41.7\% to 12.6\%. Our results show that abstention is a distinct capability that current benchmarks fail to measure.
Chat is not available.
Successful Page Load