An Information-theoretic Framework for Auditing Unfairness in Training Data
Abstract
Quantifying bias in training data---independently of any downstream model---is critical for fair machine learning, especially when data collection and model design occur in silos. Yet most fairness research focuses on model-level mitigation, leaving the data side underexplored. We address this gap by deriving multiple information-theoretic measures for auditing unfairness bias in data without requiring access to model predictions. We define our measures over feature sets and use Shapley values to deduce feature-level contributions to global unfairness bias. For some of these measures, we prove that Shapley values moderate redundancy across feature coalitions, matching empirical evidence in prior work. We develop our measures within an axiomatic framework that embeds a fairness notion, such as statistical parity (SP) or equalized odds (EO), into desired and undesired independence properties. We characterize baseline measures that guide those we develop, for example, by showing that a baseline upper bounds the bias of any downstream restricted predictor. Finally, we show that SP- and EO-aligned independence properties can conflict, paralleling Kleinberg et al.'s result on incompatible fairness notions. We validate our measures through an empirical study on real and synthetic datasets, assessing how well they predict the true bias of an extensively tuned neural network trained on the same data, and via a simple feature selection experiment. For synthetic data, we propose a parametric linear structural causal model that enables controlled generation of diverse correlation structures and bias types. Overall, our analysis provides a theoretically and empirically validated guideline for selecting an unfairness measure given a group fairness notion and testable data conditions.