Diagnosing Data Needs in Synthetic-Prior Time-Series Foundation Models
Abstract
Time-series foundation models pretrained from programmatic synthetic data have an unusual property: their training distribution is an inspectable generative program. When such a model fails on a capability, we ask whether it needs more exposure to structure already represented by its prior or genuinely new generative structure. We introduce prior-referenced diagnosis, which combines an A-score measuring under-training relative to a balanced reference model with a B-score serving as an operational proxy for coverage by existing generator families. A frozen decision rule then prescribes either additional exposure to an existing generator or new generative structure. On TempoPFN, the rule correctly diagnoses all eight planted deficits; at matched compute, prescribed continued pretraining recovers 87–131% of each planted gap, whereas continued training on the original mixture recovers little. A within-family GP restriction further localizes the failure to the removed sub-range and repairs 92–106% of that gap. We further show that a correct prescription is not automatically a safe intervention: high-learning-rate adaptation incurs a capability-specific tax, while reducing the learning rate and carrying prescribed data within the original mixture substantially limits collateral damage. Tests on a released checkpoint, a second synthetic-prior substrate, and real data delineate important boundary conditions, including pretraining-seed variability and saturation of the coverage proxy on real data. These results show that an explicit synthetic prior can serve not only as a source of pretraining data, but also as a reference for diagnosing what data a time-series foundation model needs and testing whether the resulting intervention actually repairs it.