Logically Consistent Accurate Forecasts
Abstract
Agentic forecasting systems based on large language models (LLMs) are now competitive with expert human forecasters, but these systems still exhibit a critical limitation: the inability to produce forecasts which are logically consistent. For example, forecasts for related events often fail to satisfy the axioms of probability, which in turn implies that related forecasts cannot simultaneously be correct and undermines their utility for decision makers. Enforcing consistency is not trivial -- naive approaches (eg. projection onto the set of consistent forecasts) can distort accuracy. We argue that the key to overcoming the apparent tradeoff between consistency and accuracy lies in examining the supporting reasoning for each forecast. We finetune a supervisory language model that leverages these explanations to jointly reconcile inconsistencies while improving the accuracy of related forecasts. On ForecastBench, our method improves the Brier score of a retrieval-augmented GPT-5.1 forecaster by 9.3% while reducing measures of structural inconsistency by 82-90%. We interrogate our method via a comprehensive series of baselines and ablations, and find that no alternative method matches our gains in consistency and accuracy simultaneously.