The Broken Model Reads Fairest: Head Collapse Conceals Demographic Disparity in Continually Fine-Tuned Encoders
Keira Chatwin ⋅ Miles Bliey
Abstract
It is well established that continual fine-tuning damages the classifier head more than the representation beneath it (Lanzillotta et al., 2026), but subsequently refitting a linear head on the frozen encoder recovers most of the lost accuracy. However, there has been little consideration of what that failure does to fairness evaluation, where it is both worse and harder to see. Group-fairness metrics measure rate differences between groups, and a head that has collapsed onto the most recent classes emits the same small subset of labels for nearly every input, meaning those rates barely differ. On a 25-profession class-incremental sequence over BERT-base, four of nine methods fall to the $\sim$5/25 recency ceiling and their equalized-odds gaps ($\leq 0.026$) come in below every functioning model's. Restated, an unguarded fairness table naively ranks the broken models first. The disparity was never absent though, only unmeasurable. Refitting the head on a labelled sample the size of the replay buffer each method already carries returns balanced accuracy to within 0.016 of joint training and brings the gap back with it, growing 5.3–8.0$\times$ for the collapsed methods; all nine equalized-odds gaps then converge into 0.149–0.169, with joint retraining included at 0.157. Disparity here is a property of the task and the representation rather than of the method. This creates a difficult problem: when accuracy reads as broken, disparity appears healthy, and repairing the first introduces problems with the second. Furthermore, encoder rankings depend on the readout and not on the fitting set. On a synthetic sequence built to isolate demographic drift, adding a fairness constraint does cut the accuracy-parity gap by a factor of 4.3 ($p=0.002$, 10 paired seeds), but it does not transfer in practice. No fairness-aware method, ours or FSW or GroupDRO, improves the deployed model's equalized-odds gap over ordinary replay. The remedy then must be an eligibility rule rather than a better metric: report prediction-level fairness only for models above a stated utility floor, and pair it with a readout the trained head cannot corrupt.
Chat is not available.
Successful Page Load