Kiwi: Slicing Models and Reconciling Adapters in Fault-Tolerant, Decentralized LLM Fine-Tuning
Harshil Patel ⋅ Parth Shah ⋅ Swayam Shah ⋅ Kunal Pai
Abstract
Fine-tuning even a small language model demands more memory than standard consumer hardware can provide. Distributing this workload across swarms of memory-constrained devices requires pipeline parallelism to partition large models, paired with data parallelism to scale throughput. We introduce Kiwi, a fault-tolerant system that shards a frozen base model across memory-constrained workers and trains local low-rank (LoRA) adapters strictly on each assigned partition. By synchronizing adapters peer-to-peer across pipeline replicas, Kiwi eliminates the need for any single device to host full parameter stacks. We evaluate three peer-to-peer merge strategies and show that layer slicing matches single-GPU held-out loss within seed variance, while replica averaging achieves a 43% wall-clock speedup on our 7-worker testbed with minimal held-out loss degradation ($+0.09$). This architecture withstands high-latency environments, retaining a 39% speedup over a simulated 75 ms wide-area network (WAN). Furthermore, Kiwi leverages frozen base weights to maintain mirrored warm standbys, enabling instant, zero-copy parameter recovery under mid-run worker failure.
Chat is not available.
Successful Page Load