Forgery Evidence Peaks Mid-Stack: Forensic Evidence Relay for Multimodal Forgery Detection
Abstract
Multimodal large language models (MLLMs) are increasingly used for explainable image forgery analysis, where a shared decoder is expected to predict authenticity, localize evidence, and generate natural-language explanations. Existing methods typically rely on final decoder states, leaving open whether these late, language-oriented representations preserve localized manipulation cues. We test this assumption with a token- and layer-wise probing diagnostic that uses pixel-level masks only for analysis to separate tampered-region and authentic-region vision tokens. Across three MLLM-based forgery backbones and multiple manipulation types, we identify a mid-layer forgery evidence peak: tampered-region tokens are most linearly separable at upper-middle decoder layers and become less separable afterward, whereas authentic-region tokens continue to strengthen. This region-conditioned asymmetry suggests competition between localized manipulation cues and late semantic aggregation. Motivated by this diagnostic, we propose Gated Forensic Memory (\method), a lightweight evidence relay that reads from the diagnosed peak layer and writes re-encoded vision-token states back to late vision-token states through a zero-initialized gated residual. \method keeps the LLM frozen, adds no extra vision tokens, and requires no pixel-level training supervision. Across three backbones and seven benchmarks, \method improves authenticity prediction, explanation quality, and spatial alignment without mask supervision; the best read layer matches the diagnosed peak, turning the diagnostic into a practical layer-selection criterion.