Who Gets Heard by Gossip? Correcting Participation Bias in Decentralized Language-Model Fine-Tuning
Abstract
We study participation bias in decentralized language-model fine-tuning when nodes participate at unequal rates. Our main setting is loss-adaptive participation: a node's participation probability decreases with its recent local loss, creating a feedback loop in which high-loss nodes contribute less and become further underrepresented. We compare inverse participation weighting (IPW) with two fully decentralized drift-adaptive corrections. Drift weighting (DW) scales an update by normalized drift relative to active neighbors; topological drift weighting (TDW) additionally removes a node-specific baseline induced by local mixing topology before converting excess drift to a bounded weight. Under a simple saturation model, the population TDW weight approaches IPW as the smoothing horizon grows and is conservative at finite horizons. Across synthetic transformers and LoRA fine-tuning up to 12B parameters, IPW is strongest under static participation, whereas DW and TDW are more robust under loss-adaptive participation; the topological correction is particularly important for stability at 12B.