Drift ≠ Damage: A Counterfactual Test of Encoder Adaptation in Time-Series Foundation Models
anupam mediratta
Abstract
Fine-tuning a time-series foundation model (TSFM) substantially changes its representations, and preservation methods exist to limit that change. But does the change tell us whether allowing encoder adaptation was useful? We test that directly with a frozen-encoder counterfactual: full fine-tuning (B) against encoder-frozen adaptation (D), their difference $B-D$ gated on a pretrained-capability inclusion criterion. Our headline finding is methodological: $B-D$ must be scored on held-out windows. Read on the windows model selection already uses, it flatters the frozen encoder: selection-window bias. Re-scored on a chronologically disjoint test split, four of 31 cells reverse sign (Chronos/ETTh1 $+6.8 \to -39.2$ pp), every one in the same direction. Second, within a backbone we find no reliable evidence that CKA magnitude predicts the held-out incremental value of allowing encoder-side adaptation under our evaluated protocols; across three architecturally distinct TSFMs (statistically within Moirai, directionally elsewhere), two Moirai-Large $h=96$ cells $0.004$ apart in CKA differ by $33$ pp. Third, the intervention exposes a regime within Moirai we name pretrained-capability degradation: in eight cells meeting the criterion, full fine-tuning leaves the model worse than not fine-tuning at all, in every seed, while a frozen encoder improves it. But the criterion does not predict where: pre-registered on eight new cells, it flagged all eight and only two degraded. We establish this regime's existence, not its prevalence: the third backbone contains no such cell.
Chat is not available.
Successful Page Load