Beyond Valuation Scores: Subset Selection and Portability in Time-Series Forecasting
Abstract
Current time-series data-valuation methods typically assign scalar importance scores to historical windows and retain the highest-valued examples for downstream forecasting. Yet the utility of this procedure depends not only on the scores themselves, but also on how they translate into a fixed-budget training set and whether that selection remains useful under future forecasting conditions. We study this distinction across three time-series valuation methods and three forecasting datasets. We find that reference-induced changes in selected subsets do not reliably indicate corresponding changes in future forecasting utility. Using complementary interventions on the valuation scores and score-to-set rule, we further find that poor selection can arise from weak score information, subset construction, or their interaction. These constructor effects are themselves not reliably portable across forecasting learners or deployment periods. Together, our results suggest that time-series data valuation should be evaluated end-to-end through the subsets it induces and their downstream utility, rather than from valuation scores or selection stability alone.