Limits and Potential of Score-Based Data Valuation: Redundancy, Complementarity, and Non-Monotonicity
Abstract
Score-based valuation methods, such as Shapley values and Leave-one-out (LOO), are widely used to assign value to data in modern machine learning pipelines, including for tasks such as attribution, selection, and pricing, yet it remains unclear when these scalar scores reliably guide downstream decisions. We show that their success is governed by three structural properties of the learning problem: substitutability, complementarity, and non-monotonicity. Substitutability (redundancy) can collapse pointwise credit, causing Shapley and LOO to fail even under monotone submodular valuations; bounded curvature limits this collapse and helps recover constant-factor approximation. Complementarity can break common score-based rules and greedy-style adaptive selection, though these effects diminish with sufficient coverage. Non-monotonicity implies that all non-adaptive methods, including score-based approaches, can fail, establishing a separation from adaptive algorithms. Our theoretical results, supported by empirical evidence, provide a structural view of data valuation and motivate a simple practical pipeline: deduplicate to reduce redundancy, ensure coverage to suppress complementarity, and then choose between score-based or adaptive methods based on non-monotonic effects.