Relative Energy Barriers for Copyright-Aware Language Model Adaptation
Abstract
Adapting language models on long-form corpora can improve domain behavior while also making protected passages easier to reproduce verbatim from short prefixes. We study this tension through \emph{relative energy gain}, the token-level log-density ratio between an adapted model and its pre-adaptation anchor on an exact protected continuation. This quantity separates ordinary predictability from update-induced copying pressure and accumulates over suffixes as a likelihood-ratio advantage for extractable memorization. Building on this view, we introduce Energy-Gated Copyright Regularization (EGCR), a training objective that augments maximum likelihood with a soft relative-energy budget on protected tokens. EGCR uses a differentiable gate combining local persistence of relative gain with anchor surprisal, so regularization is concentrated on distinctive spans whose probabilities rise in a memorization-like way while ordinary next-token supervision is retained elsewhere. Across BookMIA adaptation stress tests, training-prefix extraction, CopyBench transfer, and gate ablations, EGCR reduces exact-match and overlap-based reproduction while preserving held-out language-modeling and book-related utility. Token-level visualizations further show that the gate focuses on a small set of expressive positions associated with verbatim recovery. These results suggest that relative-energy barriers offer an effective and inspectable mechanism for copyright-aware language model adaptation.