Semantic-Bridge Federated Learning: Bridging CLIP Semantics for Heterogeneous FL
Leixuhai Xu ⋅ Xiangtao Zhang ⋅ Hailong Yan ⋅ Le Zhang
Abstract
Federated learning (FL) suffers from severe performance degradation under statistical heterogeneity, where label skew and domain shift induce representation drift and biased decision boundaries. While vision--language models such as CLIP provide transferable semantic priors, raw CLIP embeddings are noisy, highly correlated, and not directly aligned with the latent space of lightweight federated classifiers. We propose $\textbf{Semantic-Bridge Federated Learning}$ (SBFL), a framework that transfers pre-trained vision--language semantics into heterogeneous FL without accessing raw client data or requiring online CLIP inference. SBFL constructs text-verified semantic banks, learns a server-side semantic bridge that maps CLIP-derived features into student-compatible teacher spaces, and regularizes local training through mixed semantic prototype alignment and quality-aware semantic replay. Experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, Office-Caltech 10, and Digit5 show that SBFL consistently improves the average performance of representative FL backbones under both label-skewed and domain-shifted settings. Analysis further shows that the semantic bridge converts highly correlated CLIP priors into more separable teacher-space geometry, explaining its effectiveness in improving representation alignment and global generalization. Our code will be released upon publication.
Chat is not available.
Successful Page Load