Multimodal AI Detection In Two Words
Abstract
AI-generated text, images, audio, and video are now ubiquitous and difficult for humans to reliably identify as such, making scalable, automated detection a practical necessity. Existing detectors, however, are fragmented along modality lines---each typically relying on its own classifier head, dataset, and training pipeline---and the recent line of explainable detectors learns its rationales from human annotations whose faithfulness degrades as generators improve. We propose a unified method that addresses both issues by extending neologism learning from text generation steering to multimodal classification. We add two tokens, < REAL> and < AIGEN >, to the vocabulary of an otherwise-frozen multimodal LLM (MLLM) and train only their embedding rows (on the order of 2d parameters, fewer than 0.001% of the backbone) so that the model's next-token distribution under a paired prompt encodes the real-vs-AI-generated decision. Because all modalities are projected into a shared input embedding space, the same token pair applies across text, image, audio, and video without per-modality heads, and a pair trained on a strict subset of modalities transfers to held-out ones at test time. The trained tokens further admit free-form natural-language descriptions of what < AIGEN > has come to mean, decoded from the same frozen backbone and never supervised during training; we validate their causal faithfulness via plug-in evaluation. Across standard text, audio, image, and video benchmarks, our method matches or exceeds dedicated, modality-specific detectors and outperforms MLLM-based explainable detectors out of distribution, despite training orders of magnitude fewer parameters and using no human-annotated explanations.