AKTD: FDR Control for LLM Training Data Detection under Approximate Exchangeability
Abstract
Detecting whether specific text samples were used to train large language models is increasingly important for privacy, copyright and data auditing. Knockoff-based training data detection (KTD) offers a promising approach to controlling the false discovery rate (FDR). However, its effectiveness relies on null sign-flip symmetry induced by exchangeability between each null sample and its knockoff. This condition is difficult to satisfy in natural language settings. We study training data detection under approximate exchangeability, and demonstrate that deviations from exact exchangeability induce measurable null-tail asymmetry, which in turn leads to systematic FDR inflation under the standard knockoff+ threshold. This effect can be captured by pathwise bounds, which explain when and why KTD becomes anti-conservative. Based on this, we propose Asymmetry-adjusted KTD (AKTD), a series of procedures that correct for asymmetry by estimating the penalty term. We introduce an external null variant that uses a separate calibration pool containing only null samples to estimate the penalty term.This enables detection with FDR control without using labels from the evaluated candidate set. Experiments on WikiMIA and MIMIR using GPT and Pythia demonstrate that the asymmetry term accurately predicts the observed FDR expansion, and AKTD restores FDR control while maintaining effective detection capabilities.