$\texttt{dattri-LLM}$: A Unified and Efficient Library for Training Data Attribution at LLM Scale
Shixuan Liu ⋅ Junwei Deng ⋅ Pingbang Hu ⋅ Jiaqi Ma
Abstract
Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Many scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility. Existing libraries provide attribution workflows and gradient-processing interfaces, but differ in how they represent gradients, integrate with training pipelines, and support new methods and applications. We introduce $\texttt{dattri-LLM}$, a TDA library that addresses these challenges through a unified gradient interface. For supported layers, $\texttt{dattri-LLM}$ provides exact factorized and materialized gradient representations and uses a FLOP-aware cost model to select between them per layer and per operation. An open-loop} interface captures gradients through autograd hooks during existing backward execution, supports integration with unmodified training pipelines under DDP and FSDP, and caches gradients for reuse across compatible queries and methods. An attributor interface over a uniform gradient stream supports new methods, while training-state callbacks support applications that act on captured gradients during training. Across the evaluated attribution methods and projection regimes, $\texttt{dattri-LLM}$ achieves $1.4$-$3.1\times$ speedups over the compared libraries. We further demonstrate token-level attribution from captured gradient factors and integrate online data selection into an existing fine-tuning loop with only a few lines of user code.
Chat is not available.
Successful Page Load