Too Aligned to be Real: Detecting AI-Generated Images via Cross-modal Alignment Shift
Abstract
AI-generated image detection has become increasingly important as generative models produce highly realistic visual content. Existing detectors mainly rely on visual artifacts or visual features extracted from pretrained multimodal models, while largely overlooking the structural differences between real and generated images in the joint image--text space. In this paper, we reveal a systematic Cross-modal Alignment Shift (CAS): generated images tend to exhibit stronger and more concentrated alignment with text than semantically matched real images. We verify this phenomenon through retrieval preference, image--text residual structure, spectral concentration, and matching entropy. Motivated by this finding, we propose a CAS-based detection framework(CASNet) that learns an amplified multimodal model with a text-guided alignment-shift objective. The amplification-induced feature increment is extracted as an explicit cross-modal cue and coupled with amplified visual representations, enabling the detector to exploit both alignment-level structural evidence and visual discriminative cues. Extensive experiments across diffusion models, GANs, autoregressive models, and commercial generators demonstrate consistent improvements over state-of-the-art detectors. In particular, on recent commercial generators, our method improves ACC and AP by 5.48\% and 6.13\%, respectively.