Fold-Mean Preservation Predicts Weight Merging Success in the Cooperative Task-Vector Regime
Abstract
The model-merging toolkit was built for conflicting task vectors: constituents fine-tuned on different objectives, which merging methods exist to reconcile. A large and growing class of practical settings instead produces cooperative task vectors, whose constituents share an initialisation and an objective and differ only in the data they see. K-fold cross-validation is the limiting case, and it is the default fine-tuning protocol wherever labels are scarce: a single model must eventually be deployed, but K were trained. We ask how the existing toolkit behaves in this regime. Through 5,237 paired evaluations covering 11 merging methods, 5 pretrained molecular encoders spanning three architecture families (SMILES transformers, 2D GNNs, and a 3D equivariant network), and 10 property-prediction tasks under chemically realistic out-of-distribution splits, we find that how faithfully a merge preserves the fold-mean update, the arithmetic mean of what each constituent learned relative to the shared initialisation, predicts whether it succeeds (Spearman ρ = 0.94, Pearson r = 0.90 across methods). We formalise this as a composite three-axis score over the direction, coverage, and scale of the merged update. Ablating it shows that coverage and scale carry the signal, while direction is preserved by construction in this regime. That last property belongs to cooperative constituents specifically, and it yields a falsifiable prediction for the multi-task case. We further show why merging helps: it beats the constituent mean on 79.6% of repeats but the best constituent on only 15.2%, the pattern produced by variance reduction and not by information pooling. Genuine pooling is confined to the encoders whose folds are actually diverse. We close with a deployment default and a free preflight check, computable from the K checkpoints alone, that predicts whether merging will help at all.