Grounded Exploration for Open-Ended Image Editing
Abstract
Open-ended image-editing requests, such as ``decorate this room,'' admit multiple valid realizations. Direct sampling may produce ungrounded variation, whereas explicit diversity prompting often produces superficial differences. We study how to ground these requests without prematurely eliminating their legitimate ambiguity. We introduce a two-stage framework that first identifies diverse, feasible edit scopes and then realizes them as executable edit instructions. Across 300 open-ended image-editing tasks, our approach achieves 24.6% higher mean semantic dispersion and 17.1% higher mean groundedness than direct diversity prompting, while producing more valid and distinct realizations. These results show that structuring feasible edit directions before realization supports more useful grounded exploration.