Well-Calibrated Until It Matters: EventRisk-Bench for Conditional Reliability of Time-Series Foundation Models
Jatin Sharma ⋅ Peeyush Tapadiya
Abstract
A forecasting model can benefit from event information without producing reliable uncertainty estimates during those events. EventRisk-Bench evaluates this distinction around SEC-acceptance-aligned Item 2.02 results disclosures for U.S. equities. We forecast the log of a daily variance proxy formed from squared hourly returns; its first component may span overnight. We compare Chronos-2, TimesFM-3, and Moirai-1.1-R-Base using identical 512-session histories and controlled event inputs. Across 1,987 test events and 50,000 sampled ordinary origins, probability calibration error is $0.004$–$0.009$ when pooled over the eligible weighted population, but $0.085$–$0.394$ on events. Event conditioning reduces weighted interval score by $21.7\%$ for Chronos-2 and $53.3\%$ for TimesFM-3, yet all three have worse event scores than a rolling Event-HAR baseline with the same history. Event-specific validation recalibration reduces event calibration error in all three models; only Chronos-2 meets every prespecified repair criterion. Follow-up matching and synthetic experiments support the distinction between useful event information and reliable uncertainty. Aggregate calibration and better event scores are insufficient evidence of event reliability.
Chat is not available.
Successful Page Load