Not All Label Errors Are Equal: Efficient Task-Aware Auditing with Confidence Sequences
Fangyuan Lin ⋅ Noah Krever ⋅ Jakub Cerny ⋅ Lily Xu
Abstract
We study efficient, adaptive auditing of categorical labels in a fixed dataset, motivated by biodiversity-credit programs in which expert verification of species records is costly and downstream scores (i.e., payment) are nonlinear functions of the corrected labels. We represent label corrections with a submitted-to-verified confusion matrix. For each row (signifying a class of submitted labels), we construct a confidence sequence under randomized sampling without replacement, and aggregate these row-wise confidence sequences into a confidence sequence for the full matrix. Projecting this confidence sequence through the downstream scoring functions then yields a confidence sequence for the score. To extend this result into a practical auditing framework, we then develop a task-aware adaptive policy that prioritizes which row (e.g., species class) to inspect next based on their remaining uncertainty and their influence on the downstream score. We show that adaptive row selection and data-dependent stopping preserve simultaneous $1 - \alpha$ coverage of the downstream score, provided records are sampled randomly within each row. Together, these results provide a general framework for anytime-valid, decision-aware auditing of labels in finite datasets.
Chat is not available.
Successful Page Load