Separating Bandwidth from Windowing in Spectral-Shift Evaluation of Time-Series Foundation Models
Abstract
Spectral-shift evaluations of time-series foundation models summarize robustness with a degradation curve: accuracy against sampling rate, read as the rate at which a frozen representation stops carrying the signal. On event-localized data that curve bundles bandwidth removal with the evaluation's own windowing and reports only their sum, because a window fixed in samples rather than seconds lengthens the record as it decimates. We propose a protocol that separates the two, pairing a decimation ladder with a band-limiting ladder at fixed rate, both read through a window held fixed in seconds rather than samples, with a linear physical-feature baseline and randomly initialized twins as controls, and measure it on power-quality disturbance classification, the setting where the ground-truth event parameters the correction needs are available by construction. Under that correction, all three models still fall steeply on both ladders, so the loss survives holding duration and visibility fixed. However, once duration is held fixed, the model-dependent sign reversal an uncorrected ladder comparison reports disappears, so how the window is specified mainly decides the disagreement. A label-free statistic computable from the sampling specification alone, harmonic order retained, ranks monitoring regimes out of sample here and on a second generator, but does not size the loss to the accuracy registered in advance. A spectral-shift evaluation therefore cannot be read as spectral, or used to tell a utility which monitoring rate a feeder needs, unless it reports window placement, window duration in seconds, and event-visibility fraction at every rung.