Better Forecasts, Barely Better Rankings: Time-Series Foundation Models on Item Engagement Streams
Maksim Utushkin ⋅ Alexander D'yakonov
Abstract
Daily engagement counts of items on a media platform are sparse, bursty and short-lived series, and ranking fresh items by their future popularity is a decision platforms make every day. We audit zero-shot time-series foundation models (TSFMs) on daily listen counts of 2.8M tracks over 300 days (Yambda-500M), restricting ourselves to checkpoints released before the data (Chronos-T5, Chronos-Bolt, TimesFM-2.0, Moirai-1.1) so that pre-training contamination is impossible, and compare them with statistical baselines and an in-domain global LightGBM. On forecasting, zero-shot models beat the in-domain model in scaled point error on sparse and mature series (Moirai MASE $0.71$, TimesFM $0.74$, LightGBM $0.79$) but not in quantile loss, and every model is worse than a trivial forecaster on items younger than two weeks (MASE $2.5$--$4$). On the decision, ranking 20K fresh items by forecast engagement over the next 1, 3 or 7 days, the head of the ranking is owned by persistence: no model changes NDCG@100 by more than $+0.003$ over the last observed count at any horizon. Forecasts pay only in the tail: among items outside the previous day's top-100, zero-shot TSFMs gain $+0.015$--$0.02$ NDCG@100 and an in-domain gradient-boosted model trained on the catalogue gains $+0.03$; on the ordering of all active items the in-domain model gains $+0.03$--$0.06$ Spearman, growing with the horizon, the TSFMs $+0.01$--$0.02$. Better zero-shot forecasts buy a small, tail-only improvement that a cheap in-domain model doubles; on a second, sparser domain (short-video plays, KuaiRand-27K) the zero-shot models never beat persistence in ranking while the in-domain model does. We release the protocol as an applicability map by item age and sparsity.
Chat is not available.
Successful Page Load