Co-evolution: A "One-to-many" LLM Fine-Tuning Paradigm
Abstract
Fine-tuning has become a central mechanism for adapting large language models (LLMs) to downstream tasks, yet most existing methods follow a one-to-one paradigm: a single pretrained model is optimized into a single improved policy. This paradigm optimizes average single-policy performance, leaving no mechanism to distinguish redundant successes from genuinely complementary contributions. In this paper, we propose Co-Evolve, a one-to-many LLM fine-tuning framework that transforms a single base model into a cooperative population of persistent, specialized descendants and explicitly optimizes their complementarity. To reduce redundancy among descendants, we introduce a marginal-coverage objective that rewards each model for paying attention to examples under-covered by its peers. We instantiate this objective with evolution strategies, enabling derivative-free population-level optimization. We further introduce a seed-replay checkpoint mechanism that reduces the disk storage required to maintain multiple descendants by avoiding storing multiple full checkpoints during deployment. Experiments across multiple domains show that Co-Evolve consistently improves ensemble-level performance over strong fine-tuning and ensemble baselines. Our results suggest that optimizing for complementary model-level specialization provides a scalable alternative to single-policy fine-tuning. Our code can be found at https://anonymous.4open.science/r/Co-Evolution-Core-6604.