Evaluating Data Attribution Through Filtering Efficacy
Abstract
Data attribution explains model behaviors in terms of training data and has promising applications in model debugging and training data curation. In practice these applications often measure the filtering efficacy of removing the top <1% of documents, but the typical evaluation metric, the linear datamodeling score (LDS), is a rank correlation that does not predict how large an effect such an intervention will have, and existing scaling-law work focuses on the effects of removing 60%-90% of the training data. We formulate the query loss difference (QLD): the difference in query loss between filtering out top-k attributed documents and removing k documents selected uniformly at random. Using QLD, we study the scaling behavior of EK-FAC across finetuning datasets of 4K-512K documents, two model scales, and degrees of filtering from 40 documents to 10% of the dataset, and make comparisons with MAGIC and BM25. We find that for a given optimizer, LDS is mostly flat across dataset size and training hyperparameters, whereas QLD is sensitive: it exhibits approximate power-law scaling with degree of filtering and scales sublinearly with the percentage of data removed, while filtering efficacy at a fixed dataset size and filtering budget scales at a lower exponent as model size increases. Across dataset scales, MAGIC consistently shows greater filtering efficacy than EK-FAC, while BM25 at many points approaches EK-FAC's efficacy. We hope these results are a first step towards scaling laws for data attribution that predict a priori how much data must be removed to suppress a capability.