Beyond Correlation: Validity Conditions for Task-Affinity Evaluation
Abstract
Task-affinity scores are widely used to decide which tasks to train together and how to share model capacity. They are often validated by correlating them with observed multi-task gain. We argue that this correlation conflates four conditions that determine what it can support: the score must identify the relevant mechanism; the evaluation metric must measure the benefit of that mechanism; the candidate configurations must leave enough room for pair-specific selection to matter; and the resulting rule must generalize to the intended deployment setting. We evaluate 479 task pairs from CelebA attributes, QM9 molecular properties, and 12 multi-output regression families, covering six configurations, a three-budget experiment, and three random seeds. In each comparison, we keep the task setting fixed and vary one part of the evaluation. The four requirements can fail separately: a score can correlate with overall multi-task gain without identifying the mechanism behind it; calibration can reverse the apparent effect of encoder sharing; configurations can differ substantially while offering little benefit from pair-specific selection; and gains under random pair splits may not persist when entire task families are held out. We name these construct, metric, decision, and generalization validity. Together, they yield seven checks for establishing when an affinity score can support configuration selection.