Dynamic Model Merging Made Slim
Abstract
Model merging combines fine-tuned models into a unified one without joint training or access to original data. Dynamic merging improves flexibility by selectively activating task-relevant parameters, but existing methods either maintain a full shared model with tiny experts or allocate excessive capacity to experts, leading to suboptimal accuracy–efficiency trade-offs. We propose DiDi-Merging, a slim dynamic merging framework that uses differentiable rank allocation to balance shared and per-task expert parameters. The method casts parameter budgeting as differentiable rank optimization over low-rank modules, supervised data-free by the original task vectors, and applies a brief refinement stage to recover task fidelity. Across vision, language, and multimodal benchmarks, DiDi-Merging matches prior dynamic baselines at 1.24× the parameters of a single fine-tuned model and surpasses them at 1.4×, substantially more compact than methods requiring ≥2× storage.