FedSEM: Mitigating Cross-Client Evidence Drift in Federated Multiple Instance Learning
Abstract
Federated Multiple Instance Learning (MIFL) has emerged as a promising paradigm for privacy-preserving weakly supervised learning, particularly in medical image analysis where instance-level annotations are expensive or unavailable. Existing federated MIL methods mainly rely on parameter aggregation to learn a global model across clients. However, in MIL, the key challenge is not only client-level data heterogeneity, but also the ambiguity of instance-level evidence. Since supervision is only provided at the bag level, each client must independently infer discriminative instances through local attention mechanisms. Under heterogeneous data distributions, different clients may focus on inconsistent or even misleading instances, resulting in cross-client critical evidence drift. To address this problem, we propose FedSEM, a Federated Shared Evidence Memory framework that introduces an evidence-level communication pathway in addition to standard model aggregation. Specifically, each client identifies high-confidence evidence instances using local attention scores and compresses them into compact prototypes with importance scores. The server then refines these prototypes through redundancy removal, diversity selection, and score normalization to construct a shared evidence memory. Importantly, FedSEM does not require transmitting raw patches or dense instance embeddings; it only exchanges a small number of evidence prototypes and periodically broadcasts the refined memory, thereby introducing limited additional communication overhead compared with standard federated training. The shared memory serves as cross-client evidence anchors to guide local attention learning and instance-level representation optimization. By explicitly aligning critical evidence patterns across clients, FedSEM mitigates evidence drift while maintaining communication efficiency and privacy preservation. Extensive experiments under heterogeneous federated settings demonstrate that FedSEM consistently improves MIL performance and generalization over existing federated learning baselines.