The Calibration Trap in Federated Learning
Md Ahsan Karim
Abstract
Federated learning (FL) commonly uses sample size weighting for model aggregation. The same weighting can combine per-client expected calibration error (ECE) into client-weighted ECE (cwECE). We ask what client level reliability claims this aggregate supports. Proposition 1 shows that cwECE can approach zero while one client's ECE remains maximal, making the ratio between tail and aggregate calibration error unbounded. Proposition 2 gives a weight-dependent upper bound on the worst-client gap that becomes less restrictive as the worst-calibrated client's aggregation weight decreases. On CIFAR-10, CIFAR-100, and FEMNIST, five-seed experiments with client-specific evaluation sets drawn from the official test split yield raw-FedAvg TailTrap values of $1.60\pm0.02\times$, $1.72\pm0.16\times$, and $2.29\pm0.53\times$, respectively. The gap remains above one under three post-hoc temperature-scaling baselines, alternative ECE estimators, and a three-seed Dirichlet ablation over $\alpha\in{0.05,0.10,0.30,0.50}$. Two training-time interventions also fail to consistently reduce TailECE across five seeds. This shows that a low sample-weighted aggregate does not guarantee low calibration error for all clients. Therefore, cwECE should be reported together with a tail-sensitive client-level statistic.
Chat is not available.
Successful Page Load