FU-Author: Attribution-Based Compensation Is Incentive-Incompatible for LLMs
Abstract
We show that attribution-based compensation mechanisms proposed to pay authors for their work can perversely reward AI companies for paying them less. AI companies can retain value from authors' works while shifting attribution, and therefore compensation, toward copyright-free data of their choosing. In one illustrative attack, a model generates Star Wars-like text that TRAK attributes to Shakespeare. We show that AI companies can accomplish this without subterfuge, corruption, or alteration of authors' data. Our `attacks' on fair compensation require neither knowledge of the exact attribution mechanism nor queries to it. Nor do they require an AI company to break the proposed attribution rules. Instead, we show that proposed attribution mechanisms are incentive-incompatible, even under unrealistically idealistic assumptions. Counterintuitively, we show that mainstream attribution methods used to implement proposed compensation mechanisms (including TRAK and LoGra) are vulnerable to FU-attacks that leverage the same post-training tools used to extract more value from authors' data: (1) Fine-tuning on authors' data and/or (2) Unlearning on copyright-free data. On NanoGPT with TRAK attribution, these attacks decrease all authors' compensation by 62.5% while preserving the model's ability to generate output in the authors' style. In brief, mainstream data attribution methods designed to compensate authors can instead enable and incentivize AI companies to defund authors unless the payment mechanism is redesigned to be incentive-compatible with companies' ordinary incentives to minimize costs.