GPT-Image-Edit-1M: An Auditable Million-Scale Dataset for Instruction-Guided Image Editing
Abstract
We present GPT-IMAGE-EDIT-1M, a million-scale instruction-guided image editing dataset in which every sample includes an auditable curation decision and a traceable instruction trail. Open training data for image editing often suffers from instruction-image mismatch, weak edits, and opaque filtering; simply scaling raw data does not address these issues. To improve data quality, we regenerate edited outputs using GPT-Image-1, augment samples with complex multi-step edit instructions, and apply a policy-based quality control pipeline that assigns KEEP/RELABEL/DROP decisions with reason tags. RELABEL samples receive aligned instructions that better match the observed edit, and a second judge independently re-runs the same protocol to measure decision stability. After quality control, average quality improves from 6.90 to 7.67 under a shared rubric. Across two frozen encoder configurations, T5 and QwenVL+T5, models fine-tuned on GPT-IMAGE-EDIT-1M outperform the public FLUX.1-Kontext-dev baseline on key benchmarks, including GEdit-EN (7.22 vs. 6.26), ImgEdit (3.93 vs. 3.52), and KRISBench (66.23 vs. 49.54), while remaining competitive on Complex-Edit and OmniContext SINGLE. Controlled data-swap experiments show that GPT-Image-1 regeneration provides the dominant downstream gain, while quality control yields a smaller but consistent improvement over a size-matched random regenerated subset. Concatenating frozen QwenVL embeddings with T5 substantially improves inferential-complex benchmarks such as RISEBench and KRISBench without degrading general-editing performance. Finally, as an auxiliary diagnostic rather than a primary ranking protocol, inference-time instruction rewriting raises GEdit-EN to 7.75 and KRISBench to 78.55, suggesting that prompt specificity exposes additional instruction-following capacity in the fine-tuned editors.