When More Capacity Hurts: Unequal Parameter Exposure in Model-Heterogeneous Federated Learning
Abstract
Model-heterogeneous federated learning trains nested submodels: all clients update a shared core, while only a capable subset updates width- or depth-exclusive parameters. Consequently, the training exposure of each parameter group depends on which clients can host it. We show that, under sparse full-model eligibility, this asymmetry can reverse the preferred deployment: on CIFAR-100, a calibrated quarter-width core outperforms its full model by 2.7 percentage points at 10\% eligibility, whereas the full model leads by 6.1 points at 60\%. The crossover replicates on CIFAR-10 and a width-sliced Vision Transformer, while the same sparse-exposure harm extends to depth-heterogeneous models. Controlled interventions rule out label skew, writer count, and federated aggregation as sufficient explanations, while a matched-optimization experiment separates the effects of unique-data volume and cumulative sample presentations. Holding optimization dose fixed, shrinking the shell's unique-data pool leaves the deployment gap essentially unchanged. A large deficit emerges when both data diversity and optimization dose are restricted, supporting an interaction between these two components of parameter exposure. Under sparse exposure, a matched homogeneous small model also performs comparably to the strongest calibrated extracted core. The same exposure asymmetry can arise in decentralized foundation-model training whenever heterogeneous compute restricts which participants can update particular parameter blocks. These findings motivate exposure-aware evaluation that compares full, extracted, and homogeneous deployments across eligibility levels.