Efficient Label Distribution Inference Attack and Defense on Classifier Weights in Federated Learning
Abstract
Label distribution leakage in federated learning (FL) poses a critical privacy threat, exposing population-level patterns beyond individual data breaches. Building on known correlations between class sample frequencies and classifier weight magnitudes, we demonstrate how adversaries can exploit these relationships to infer label distributions from shared model parameters in FL. To characterize this threat, we propose WLIA (Weight Norm-based Label Distribution Inference Attack), which infers client label distributions by analyzing variations in classifier weight norms across communication rounds. WLIA operates on the classifier layer without requiring auxiliary data, synthetic samples, or meta-classifiers, enabling broader applicability across heterogeneous model architectures. To counter this threat, we introduce RNR (Randomized Weight Norm Regularization), a defense mechanism that disrupts frequency-weight correlations by randomly regularizing weight norms during local training. RNR adds only a penalty term to the cross-entropy loss, making it efficient and easy to integrate. Comprehensive experiments across six datasets demonstrate that WLIA achieves superior inference accuracy compared to existing attacks, while RNR effectively mitigates leakage with minimal impact on utility.