Elicited Adaptation: Auditable Localized Fairness via Pairwise Queries
Shrey Shah
Abstract
A single global fairness criterion cannot serve communities whose conceptions of fair treatment differ. We introduce \emph{elicited localized fairness}: an end-to-end pipeline that elicits per-community Mahalanobis metric--tolerance specifications $\cF_k = (d_k, \eps_k)$ from stakeholders via two phases of pairwise queries (similarity, then tolerance acceptability), trains a localized classifier (\textsc{LocAdapt}), and audits the deployed model on a held-out tolerance fold. \emph{Theory.} A constrained-MLE rate $\tilde O(\sqrt{d^2/n})$ for metric elicitation and a matching $\Omega(d^2/t^2)$ active-query lower bound, exact in $(d, n, t)$ exponents; a held-out audit certificate that converts Phase 2 queries into a PAC-style violation guarantee. \emph{Empirics.} On heterogeneous-community settings (Stress $K{=}5$, Adult-K3-human; 10 seeds), \textsc{LocAdapt} sits on the (accuracy, strict $d{\leq}1$ violation) Pareto frontier: averaged and matched-accuracy Global IF lose on strict at comparable accuracy ($40$--$80\%$ higher), and pooled-MLE Global IF matches strict only at $-2$pp accuracy. A five-axis diagnostic (acc, strict, score-diff spread, prediction entropy, AUC) confirms the gain is not prediction-smoothing. Plugging Phase 2 elicited tolerances into \textsc{LocAdapt} training reduces the COMPAS deployed-spec violation $66\%$ ($p{\ll}10^{-4}$). Five pre-registered human studies on COMPAS (360 annotators total) validate the pipeline end-to-end; a 180-annotator scorer-stability ablation across three scorer architectures preserves the cross-community $\hat\eps$ ordering and pools to $p{=}0.0014$ ($n{=}90$ vs.\ $90$).
Chat is not available.
Successful Page Load