Passing the Accuracy SLA While Failing the Fairness One: Disparity Drift and Its Cost in Continually Updated Classifiers
Abstract
Enterprise classifiers are retrained as data arrives and are typically governed by service-level objectives written in aggregate accuracy, holding performance on what the system already handles to within some tolerance. However, focusing solely on that one metric can miss how well the system performs across different groups. In this paper, we show that a deployment can satisfy that objective indefinitely while its performance for a demographic group degrades. We then simulate an environment that an operator tracks to understand model performance (mean retention on already-handled tasks, alarming past a 10% drop). For the strongest replay method the utility alarm remains silent on all 10 seeds, while a parallel disparity alarm would fire on all 10—evidencing the need for both metrics. On a 25-profession sequence built from professional biographies, a résumé-screening workload with an existing compliance regime, we compare nine methods under a matched memory budget. Two gradient-projection constraints cut the accuracy-parity gap, the spread in per-group accuracy, by factors of 4.1 and 4.3, at 1.75× and 1.44× their backbones' training cost—an overhead that is fixed per step and stable across environments. The closest prior method, FSW, is considerably harder to budget for. Its per-epoch linear program dominates its runtime, and running the identical configuration under the same pinned dependency versions in two environments produced 96× and 0.79× ordinary replay, a spread of ∼122 for statistically indistinguishable results, while every other method moved by less than 1.1×. Notably, an intervention whose cost is set by a solver stack cannot be capacity-planned from a published number. We then stress-test these methods on the multiclass sequence, where the failure reverses entirely, in that the utility alarm fires while the disparity alarm goes quiet. Here four of nine methods degenerate onto the five newest of 25 professions, and such a classifier has almost no measurable disparity, so its equalized-odds gap (≤ 0.026, against 0.146 for a system that still works) comes in under every working model's. Refitting the head returns the gap, growing it 5.3–8.0×, so a quiet disparity alarm is evidence of fairness only while the utility alarm is quiet too—meaning that remediating an accuracy incident will look like it caused a fairness one. Mitigation also stops working there, since no fairness-aware method improves deployed-head equalized odds over ordinary replay, and against its own backbone Fair-DER++ is worse (p = 0.010, 0.078 corrected). Still, we conclude that the limited results of fairness mitigation reinforce the need to measure fairness itself, disaggregated, continuously, and gated on a utility precondition, rather than inferred from the aggregate accuracy the objective was written to preserve.