TS-AgentBench: Necessity vs. Robustness in Agentic Forecasting
Malik Tiomoko ⋅ Youssef Attia El Hili ⋅ Bahaeddine Abdessalem ⋅ Ding Wang ⋅ DONG Zhiwei ⋅ Zehao Xiao ⋅ Ambroise Odonnat ⋅ shifeng xie ⋅ Lei Zan ⋅ Jiale Zheng ⋅ Tareq Si Salem ⋅ Zhang Keli ⋅ Jianfeng Zhang ⋅ Lujia Pan
Abstract
LLM agents for time-series forecasting are increasingly evaluated on end-to-end benchmarks, but current evaluations cannot tell whether an agent is robust to feature corruption or simply ignores the feature. We introduce \textbf{TS-AgentBench}, a dual-cohort benchmark combining real-world forecasting tasks with a structural causal synthetic environment that enables controlled analysis of agent behavior. The benchmark evaluates agents under three conditions: clean input, corrupted covariates, and missing covariates, with a bootstrap-based necessity test to measure whether features are actually used. We find that some agents rely on misleading covariates and can perform better when these covariates are removed. Across real and synthetic cohorts, agent rankings remain consistent ($\rho = 0.786$, $p < 0.01$), while our analysis reveals important differences in how agents exploit available information.
Chat is not available.
Successful Page Load