Efficient and Accurate Evaluation of Training Data Attribution
Simon Vary ⋅ Alireza Mousavi-Hosseini ⋅ Roger Grosse
Abstract
Evaluating training data attribution (TDA) methods is costly, since evaluation involves retraining a model on many different subsets of the original training set. For example, the Linear Datamodeling Score (LDS), which has become a standard metric for evaluating TDA methods, measures the correlation between TDA predictions and retrained model outputs. In this paper, we reduce the amount of computation needed to estimate the LDS along three axes: the number of random retraining subsets (masks), the number of random seeds for each retraining subset, and the ensemble size for estimating attribution scores of averaged TDA ensembles. Each of these factors induces a form of bias in the LDS estimation, and removing these sources of bias significantly reduces the number of retrainings needed to match the same accuracy in LDS estimation. On ResNet-9/CIFAR-10, finite-mask debiasing achieves $3.2\%$ relative RMSE with as few as $4$ masks, and $32$ masks match the accuracy of ordinary Spearman using $96$ masks, measured against a 300-mask reference. On ResNet-9/CIFAR-10 and BERT/QNLI, correcting the training noise reaches the same accuracy with $2\times$ to $5\times$ fewer retrainings.
Chat is not available.
Successful Page Load