Imperfect Influence, Reliable Rankings: A Theory of TRAK for Data Attribution
Abstract
Data attribution, which traces a model’s prediction back to specific training data, is an important tool for interpreting sophisticated AI models. The widely used TRAK algorithm addresses this challenge by first approximating the underlying model with a kernel machine and then leveraging techniques developed for approximating the leave-one-out (ALO) risk. Despite its strong empirical performance, the theoretical conditions under which TRAK approximations are accurate, as well as the regimes in which they break down, remain largely unexplored. In this paper, we provide a theoretical analysis of the TRAK algorithm, characterizing its performance and quantifying the errors introduced by the approximations on which the method relies. We show that although the approximations can incur significant errors, TRAK preserves the separation between highly influential and weakly influential data points. We corroborate our theoretical results through extensive simulations and empirical studies.