Do LLM Forecasts Change Their Minds Coherently?
Adarsh Agrawal
Abstract
Forecasting benchmarks usually evaluate where a model finishes rather than how its predictions evolve. We introduce a martingale audit for the full probability path. For a coherent forecast, expected squared probability movement equals expected uncertainty reduction; their difference yields an intermediate-label-free residual, $C=M-R$. Across 73 resolved markets, an eight-model panel exhibits a model-crowd coherence gap of $+0.061$. This difference remains under alternative prompts, observation grids, missing-path treatments, and a deliberately noisy repeated-elicitation stress test. It also reappears in a prospectively collected set of 41 resolved markets. The gap is largest among efficient systems, while the strongest evaluated systems remain close to the crowd. Because models and market participants have different information sets, our results support a comparative rather than causal interpretation. Nevertheless, endpoint scores alone miss a systematic difference in how LLM forecasts change as evidence arrives.
Chat is not available.
Successful Page Load