A Ceiling on Query-Independent Data Attribution
Ahmet Erdem Pamuk
Abstract
Training-data attribution assigns each training example a score for each test query, but some common evaluations reward only the query-independent component of those scores. We study this mismatch in a convex setting where exact leave-one-out retraining is feasible for all 2,000 training examples and 200 test queries. We decompose each attribution matrix as $\tau(i,q)=a(i)+b(q)+r(i,q)$, where $a(i)$ is the per-example main effect, $b(q)$ is a per-query offset, and $r(i,q)$ is the interaction. This lets us compute the best possible query-independent attribution score directly from the exact ground truth. Its ceiling is only $\rho=0.0873$, close to the analytic bound of $0.0936$, while the ground-truth query-independent variance share is $0.0050$, essentially at the within-query permutation floor of $0.0049$. We show that this small main effect arises because test gradients cancel across queries, causing the query-independent share to scale approximately as $1/M$. Despite this, query-blind baselines can perform nearly as well as exact influence functions on mislabel detection: AUROC $0.965$ versus $0.966$. On query-specific counterfactual removal, however, the same baseline scores $0.000$ while exact influence functions reach $0.920$. We also find that attribution conclusions are highly sensitive to configuration: matching TRAK’s ridge to the training regularizer raises its rank correlation from $0.3535$ to $0.8159$, and evenly spaced TracIn checkpoints collapse to a rescaled gradient dot product in this convergent convex setting. These results show that attribution methods are not merely repackaged difficulty scores; rather, some standard evaluations fail to distinguish genuine query-specific attribution from query-independent rankings.
Chat is not available.
Successful Page Load