Unified Forensic Preference Learning for Generalizable Synthetic Image Detection
Abstract
Despite rapid progress in synthetic image detection, detectors that perform well on known generators often become brittle when the source model, post-processing pipeline, or scene composition changes. A key reason is that standard binary supervision teaches models to separate training distributions rather than to identify authenticity-relevant evidence: low-level artifacts become shortcuts, while higher-level semantic implausibilities remain weakly grounded. To address this limitation, we propose UniFPL, a unified forensics preference learning framework that converts multimodal large language model (MLLM) forensic knowledge into image-specific preference supervision. UniFPL constructs forensic preference data from both controlled artifact synthesis and observed artifact harvesting, then learns from list-level evidence rankings that encode which visual-textual cues are more diagnostic of synthetic manipulation. On top of a CLIP-based discriminative detector, UniFPL jointly optimizes authenticity classification, artifacts-aware preference alignment, and semantic structure regularization, allowing the model to inherit contextual forensic reasoning while retaining the stable and efficient inference of a discriminative encoder. Across eleven curated and in-the-wild benchmarks, UniFPL achieves an average balanced accuracy of 93.4\% and improves the worst-case accuracy from 81.4\% to 86.1\% over the strongest baseline, demonstrating stronger cross-generator generalization and robustness to real-world image degradations.