Federated Dataset Simulation: Inducing Label-Free Heterogeneity Across Tasks
Vasilis Siomos ⋅ Lam Ngo ⋅ Jonathan Passerat-Palmbach ⋅ Giacomo Tarroni
Abstract
Federated Learning enables collaborative model training across decentralized clients without sharing raw data. While the field has seen rapid growth, the datasets underpinning FL evaluation have not kept pace: real federated datasets remain scarce, and most studies rely on small-scale centralized classification datasets partitioned via Dirichlet label skew, a one-dimensional proxy for the multi-dimensional heterogeneity present in real federations. Moreover, label-based partitioning is by construction inapplicable to non-classification tasks such as segmentation and detection, for which no principled simulation method exists. We cast the conversion of centralized datasets into federated benchmarks as the task of Federated Dataset Simulation (FDS), and decompose it into a representation function $\Phi$ and an assignment mechanism $\pi$. We provide two label-free instantiations---geometric partitioning via pretrained model embeddings and semantic partitioning via vision-language model-derived criteria---applicable to classification and dense prediction tasks alike. We evaluate across four datasets spanning classification, segmentation, and detection, and show that both methods induce heterogeneity that matches or exceeds that of label-skew partitioning, with measurable impact on downstream FL performance. We release \texttt{fds}, a modular open-source package for end-to-end federated dataset simulation, heterogeneity measurement, and FL training, at \url{https://anonymous.4open.science/r/FDS-neurips/}.
Chat is not available.
Successful Page Load