MarkTune: Improving the Quality-Detectability Trade-off in Model-Embedded LLM Watermarking
Abstract
Language model watermarking schemes fall into two broad categories: inference-time methods, which modify the decoding process at generation time, and model-embedded methods, which encode a secret signal directly into the model weights. Inference-time watermarking methods incur additional inference overhead and are not applicable in the increasingly prevalent open-weight model settings. In contrast, model-embedded watermarks introduce no additional inference latency---generation uses the standard sampling pipeline---and are particularly well-suited to open-weight settings. However, existing model-embedded approaches such as GaussMark face a fundamental quality-detectability trade-off: achieving strong detection power typically requires weight perturbations that noticeably degrade generation quality. We introduce MarkTune, a theoretically grounded on-policy fine-tuning framework that treats the GaussMark detection statistic as a reward while explicitly regularizing for text quality. Empirically, MarkTune substantially improves the quality-detectability frontier of GaussMark, approaching the detectability of strong inference-time schemes while preserving generation quality and downstream task performance. MarkTune is also extremely robust to paraphrasing and fine-tuning attacks, and generalizes across datasets: models fine-tuned on one corpus retain substantial detection power on unseen data.