What Does Fifty Days Stabilize? A Commit-Level Audit of ForecastBench
Daniel Alami
Abstract
A live benchmark must publish before every outcome is known. ForecastBench waits 50 days before adding a new model to its public leaderboard, after broad rank correlations appear to stabilize. A stable ordering, however, can still produce a different winner. We audit that gap by reconstructing the resolution evidence available 30, 50, and 90 days after 13 ForecastBench releases and scoring a fixed cohort of 38--54 configurations on exact common support. At day 50, median Kendall agreement with the final ordering is $\tau_b=0.824$, yet 6 of 13 winners change; three day-50 choices later fall outside the final top five, with maximum Brier regret of 0.0201. In a common 40-day follow-up across 15 releases, five day-50 winners change by day 90 and two fall outside the day-90 top five. A fixed-cohort replay of 28 configurations on 1,055 dated LiveCodeBench problems finds five changed winners across 13 cumulative states, falling to one of nine after its recommended date cutoff; every early winner remains in the final top five. Matched random supports reproduce the ForecastBench winner turnover, but the actual day-50 ordering is weaker than its release-specific matched median in 12 of 13 releases. Comparator-time and population-weighting audits identify two further ways that recorded evaluation state changes a decision. We conclude with a minimal reporting protocol for replayable leaderboard claims. The companion artifact implements it by linking every result to a commit, cohort, support set, and state record.
Chat is not available.
Successful Page Load