FedTrace: Generated-Content-Based Watermark Verification for Traitor Tracing in Federated Learning
Abstract
As large generative models become widely deployed and customized, federated learning is increasingly used to adapt them while keeping user data local. Because clients repeatedly receive up-to-date adapters, a malicious client can copy a dispatched adapter, use it offline, and monetize generated images without exposing the stolen weights or a queryable service. Existing federated watermarking and traitor-tracing methods usually assume white-box access to suspect weights or black-box query access to the deployed model. In realistic generative-model theft, however, the defender may only observe images that have already circulated. We propose \emph{FedTrace}, a generated-content-based watermark verification framework that attributes leaked federated models from suspicious generated images alone. FedTrace couples three designs: a round-wise watermark lifecycle that separates client identity distribution from global utility aggregation, a low-drift reliable-bit carrier that embeds identity in watermark positions stable under local adaptation, and anti-collision coding with soft subset verification for distinguishing singleton and collusive leakage. These components enable post-local, collusion-aware attribution from generated outputs alone, without suspect weights or online query interfaces. Extensive experiments across customized diffusion datasets, base models, and federated settings show that FedTrace preserves detectable client identities under local adaptation and strengthens subset-aware tracing against collusive leakage.