IND-Eval: Aligning Evaluation and Simulation for Industrial Time Series Foundation Models
Abstract
We introduce IND-Eval, a public benchmark that organizes industrial time series across domains into source-specific forecasting tasks. Relative to established benchmarks such as GIFT-Eval, it emphasizes finer sampling resolutions and a greater prevalence of multivariate forecasting. Evaluating Chronos-2, TiRex-2, and TimesFM-3, we find that TimesFM-3 achieves the lowest source-balanced aggregate nMASE and nCRPS and leads seven of nine sources under each metric; source-specific exceptions reinforce the need for domain-representative evaluation for model selection. These findings also motivate improving domain alignment during pretraining, where real industrial data are often too scarce to close domain-specific gaps directly; we therefore investigate whether a domain's statistical structure can guide simulator design. In a case study on two offshore well-monitoring regimes, we improve a general simulator's coverage using an oscillator-augmented simulation pipeline. This lowers forecasting error in the target domain but reduces performance on the subset of GIFT-Eval with natively univariate targets. A 50/50 mixture of the general and augmented simulation recipes achieves the lowest error among the three training recipes for all four event--metric pairs in the case study while recovering most of the lost GIFT-Eval performance. Together, these results illustrate how domain-representative evaluation can guide model selection and the composition of simulated data while revealing a trade-off between specialization and generality.