ProbTraj: Benchmarking and Post-Training for Probabilistic Forecasting
Abstract
We present ProbTraj for benchmarking and post-training probabilistic forecasters. ProbTraj-Bench pairs precisely defined events with forecast cutoffs, frozen evidence, and resolved outcomes, evaluating probability quality alongside valid-output coverage. A complementary trajectory-based method distills tool use and evidence analysis from a separate teacher corpus, then optimizes forecasts with an outcome-dominated reward. On the benchmark’s 586-case market-blind development set, the final ProbTraj-SFT and ProbTraj-RL models achieve accuracies of 68.38% and 72.12%, with Brier scores of 0.2037 and 0.1453, respectively. The reported SFT-to-RL comparison shows a 3.74-percentage-point accuracy gain and a 28.67% relative reduction in Brier score. Together, our benchmark and post-training method provide a unified framework for developing and evaluating AI probabilistic forecasting systems.