Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
Kai Yao
Abstract
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. Experiments on same-prompt real/diffusion benchmarks support this separation. Public CLIP-based interfaces collapse under targeted attacks, and stronger clean or adversarial retraining does not restore positive operational separation. In our evaluated settings, positive empirical upper bounds on the robust gap appear only under private low-bandwidth binary interfaces with abstention. The main lesson is simple: robust passive provenance requires both a source--target statistical analysis and an interface-aware evaluation of what the verifier reveals.
Chat is not available.
Successful Page Load