Heterogeneity-aware Distillation for Federated Continual Learning
Abstract
Federated Continual Learning (FCL) requires distributed clients to acquire new tasks while retaining previously learned knowledge. In this paper, we revisit the heterogeneity challenge in FCL, which is composed of two distinct aspects: spatial static heterogeneity, e.g., initial data distribution and model architecture differences, and temporal dynamic heterogeneity, e.g., task shifts over time. While existing methods acknowledge the presence of heterogeneity, they typically treat it in a unified manner by employing a single adaptive strategy without explicitly disentangling their impacts. In contrast, we propose a heterogeneity-aware distillation framework for FCL called Ha-FCL, which is designed to handle spatial and temporal heterogeneity through tailored distillation strategies. Specifically, Ha-FCL combines width-aware federated training with server-side public-anchor distillation, integrating multi-layer feature distillation for architecture-aware alignment with historical-logit regularization and representation co-distillation. This forms a coordinated dual-dimensional distillation scheme that balances adaptation to new tasks with retention of prior knowledge. This enables more targeted knowledge transfer and leads to improved model performance under complex FCL scenarios. We evaluate our method in various heterogeneous settings. Compared with state-of-the-art baselines, Ha-FCL consistently improves test accuracy and shows better robustness under challenging heterogeneous FCL settings.