quanda: An Interpretability Toolkit for Training Data Attribution Evaluation
Abstract
Training data attribution (TDA) methods estimate the influence of individual training samples on a model's predictions, offering a principled lens for interpreting neural networks. Despite rapid methodological progress, TDA evaluation remains scattered across papers and rarely reproducible, making fair comparison between methods challenging. We introduce quanda, a Python toolkit that standardizes TDA evaluation. quanda features a comprehensive collection of evaluation metrics and ready-to-use benchmarks spanning image classification, text classification, and language modeling, alongside a uniform interface for integrating diverse TDA implementations. We showcase quanda through a comparative study of existing TDA methods and find that no approach excels uniformly across metrics, with clear trade-offs between ground-truth faithfulness and downstream task performance. By consolidating the tooling around TDA, quanda equips the community to evaluate attribution methods rigorously and reproducibly. The toolkit is thoroughly tested, documented, and available as an open-source library on PyPI and at https://anonymous.4open.science/r/quanda.