ReplayMark: Probe-and-Replay Model-Response Watermarking for Diffusion Language Models
Abstract
Language models output a probability distribution over tokens, and watermarks are embedded by altering how tokens are sampled. Existing token-level schemes presuppose a fixed decoding order, whereas masked diffusion language models (DLMs) decode in parallel by confidence; imposing an order alters the decoder being watermarked. We observe that, for a fixed checkpoint, re-masking part of the decoded context raises or lowers the probabilities of other tokens consistently across forward passes, a basis for watermarking DLMs. We therefore propose ReplayMark. First, the provider compares forward passes that re-mask different parts of the context to identify positions where some tokens gain probability and others lose it, then lets the key select a direction, and resamples under the unmodified decoding schedule, so each token remains a sample from the model. Afterwards, the verifier re-masks the document, reruns the same checkpoint, and counts the positions whose shift agrees with the key. Because positions are selected independently of the key, text generated without it agrees with probability one half, so the count is exactly binomial, fixing the false-positive rate without calibration. With one configuration selected on LLaDA and reused on Dream, ReplayMark attains true-positive rates of 0.80 and 0.90 at a 1% false-positive rate, with task accuracy within 0.10 of the baseline. Verification takes 144 forward passes per 512-token document, and detection degrades under dispersed edits but not contiguous ones. Code is available at https://github.com/ming053l/ReplayMark.