Beyond Scalar Data Influence: Balancing Quality and Coverage for Fine-Tuning Data Selection
Abstract
Influence-based data selection often reduces each training example’s effects on a validation set to a single average score, enabling top-K selection by predicted utility. We study what information this scalar aggregation discards and how scalar ranking compares with set-level coverage of train-to-validation influence profiles under a fixed data budget. Across nine model–dataset settings, non-scalar components account for the majority of pairwise variation among training examples and induce highly stable geometry across validation splits. Yet emphasizing representation coverage alone does not consistently improve downstream fine-tuning. Progressively restricting coverage-based selection to examples with favorable scalar influence reveals a systematic tradeoff between marginal influence quality and geometric representativeness, with downstream outcomes varying across models and datasets. These results show that pairwise influence contains substantial, reproducible structure beyond scalar scores, while the downstream utility of that structure depends on how individual quality and set-level coverage are balanced.