The road reaches every place, the short cut only one: Self-Adversarial Shortcut Mitigation for AI-Generated Image Detection
Abstract
With the rapid advancement of generative models, AI-generated images (AIGIs) have become increasingly realistic, posing new challenges for reliable detection. In this work, we revisit the root causes of poor generalization in AIGI detectors from the perspective of shortcut learning. We find that detectors often rely on spurious semantic cues or a single prominent artifact, causing the feature space to collapse toward low diversity and severely limiting generalization. To characterize this issue, we propose the Generative Artifact Manifold Clustering (GAMC) hypothesis, which shows that although generative artifacts are multi-dimensional, different generators occupy distinct distributions along these dimensions, making detectors that rely on local artifact cues inherently fragile. To address this, we shift AIGI detection from “single-cue reliance” to “multi-dimensional discrimination” and introduce the Semantic–Artifact Self-Adversarial Feature Learning (SA-AFL) framework. SA-AFL progressively disentangles semantics from artifacts and employs a self-adversarial mechanism to balance learning across artifact dimensions, enabling the model to retain richer and more shared generative features. As a result, SA-AFL reduces shortcut reliance and learns more transferable evidence for cross-generator detection. Comprehensive experiments on images from 23 generative models show that SA-AFL improves ACC by 6.84\% and mAP by 8.25\% over state-of-the-art methods.