Evaluating Deception, Not Just Deepfakes
Abstract
Deepfake detection is evaluated as a media classification task: given a clip, decide whether it is synthetic. Detectors report high accuracy on curated benchmarks, yet these results are increasingly used to support a different claim, that deployed systems will prevent deepfake-enabled deception in video calls, voice calls, and targeted communications. We argue that this inference requires two validity steps that current benchmarks do not establish. External validity asks whether synthetic-origin classification transfers across generators, processing pipelines, and transmission channels; we show it rests on five forensic premises, all of which are eroding. Consequential validity asks whether correct classification at deployment is sufficient for prevention. Correct classification alone is insufficient because prevention also depends on context, verification, intervention timing, and recipient response. Even a perfect detector leaves this second gap open. We therefore propose evaluating the interaction scenario rather than the isolated clip, define interactive deception through three interaction-level indicators drawn from speech act theory, Gricean pragmatics, and Cialdini’s principles of influence, and specify decision-aware metrics: attack prevention rate, benign pass-through rate, precision at fixed prevention under stated base rates, and intervention latency.