A Retained-Signal Interface for LLM Watermark Robustness under Paraphrase
Chen Wang ⋅ Yinxuan Huang ⋅ Yexin Cui cui ⋅ Maoqing Zhong ⋅ LaiLong Luo
Abstract
Watermark robustness under paraphrase is usually reported as AUC against a named rewriter, making results hard to compare across watermark schemes, paraphrasers, and detectors. We propose a retained-signal reporting interface: if a watermark injects KL signal $\Delta=\KL(P_1\Vert P_0)$ and a paraphrase channel retains KL fraction $\rho$, then any detector's post-paraphrase advantage is controlled by $\Delta\rho$. We give explicit Pinsker and Bretagnolle--Huber threshold laws, prove the $\Theta(\sqrt{\Delta\rho})$ rate is sharp for any $(\Delta,\rho)$-only converse, and show the bound is implementation-blind once watermark families are matched by $\Delta$. For real LLM paraphrases, we instantiate the interface as a score-projected audit that estimates $\widehat{\Delta}\widehat{\rho}$ for a deployed detector and reports it with semantic preservation and the detector projection. In a KGW/Unigram $\times$ Qwen/Llama benchmark, this score-effective signal predicts post-paraphrase AUC across lengths, strengths, and paraphraser families. The resulting protocol separates injected signal, retained signal, semantic quality, and detector choice, replacing one-off AUC tables with a reproducible robustness report.
Chat is not available.
Successful Page Load