Watermark Robustness and Radioactivity May Be at Odds in Federated Learning
Leixu Huang ⋅ Zedian Shao ⋅ Teodora Baluta
Abstract
Federated learning (FL) enables fine-tuning large language models (LLMs) across distributed data sources. As these sources increasingly include LLM-generated text, provenance tracking becomes essential for accountability and transparency. We adapt LLM watermarking for FL data provenance, where a subset of clients computes local updates on watermarked data, and the server averages all updates into the global LLM. We find that LLM watermarks remain $\mathit{radioactive}$ in this setup: the watermark signal stays detectable after fine-tuning, with a $p$-value reaching $10^{-24}$ even when only 6.6\% of data is watermarked. However, the server can act as an $\mathit{active}$ $\mathit{adversary}$, aiming to preserve model utility while evading provenance tracking. We observe that updates from watermarked synthetic data appear as outliers relative to clean updates. Our adversary thus applies strong robust aggregation to filter out these outliers along with the watermark signal. All evaluated watermarks are $\mathit{not}$ $\mathit{robust}$ against such an adversarial server. Our work suggests fundamental trade-offs between radioactivity, robustness, and utility.
Chat is not available.
Successful Page Load