Anytime-Valid Pairwise Testing for Incrementally Revealed LLM Benchmarks
Anany Kotawala
Abstract
When paired benchmark items are revealed incrementally and a gap may be reported after interim results have been inspected, the reporting sample size is a stopping time chosen after seeing the data, and a fixed-horizon McNemar test does not control Type-I error at that time. We study the cost of replacing it with a beta-binomial mixture e-process, applied to paired outcome matrices from two archived leaderboard snapshots. Across 49 real pairwise comparisons, 37 are significant under a fixed-horizon exact McNemar test but only 33 meet the anytime-valid threshold. Because $\min(1, 1/E_N)$ is itself a p-value that direction is automatic; the count measures the price of the time-uniform guarantee in the close-adjacency regime. One adjacent-rank gap with fixed-n $p = 0.008$ has a confidence sequence, conditional on the discordance process as McNemar's test is, that still contains zero after all 12,031 items. Intervals widen by 1.6-1.9x and the items needed to resolve a gap roughly double.
Chat is not available.
Successful Page Load