MLLM-Edit: Benchmarking Image Forgery Detection and Localization under MLLM-based Editing
Abstract
Multimodal large language models (MLLMs) enable prompt-driven image editing, introducing a new challenge for image forgery detection and localization (IFDL). Unlike traditional or mask-guided editing, MLLM-based editing may re-synthesize both tampered and non-tampered regions, weakening the correspondence between tampering regions and pixel-level forensic discrepancies. To study this emerging paradigm, we introduce MLLM-Edit, a large-scale benchmark containing approximately 250k tampered images with pixel-level tampering annotations. MLLM-Edit is constructed through a structured pipeline that combines diverse real-image collection, human-like prompt design, multi-source MLLM-based editing, and semi-automatic mask annotation. Specifically, we collect natural and text images from 30 sources, formulate prompts with target content, tampering type, and consistency constraints, generate forgeries using 18 representative open-source and closed-source MLLM-based editors, and obtain pixel-level tampering masks through a semi-automatic annotation pipeline. We further establish four evaluation protocols covering traditional editing, mask-guided editing, fully synthetic generation, and MLLM-based editing. Experiments show that existing IFDL methods degrade substantially on MLLM-based forgeries, especially in pixel-level localization, highlighting the need for dedicated benchmarks and methods. A demo subset of MLLM-Edit is available at the \href{https://kaggle.com/datasets/8441a4e0594aa02aa7a59ee3eead454410d34716d14423d56a56c7bbbac80655}{project page}.