CruxBench: A Benchmark of Information Discovery
Lina Piao ⋅ Amelia Hui Dai ⋅ Nick Merrill ⋅ Nadja Flechner ⋅ Ezra Karger ⋅ Haifeng Xu
Abstract
Benchmarks for large language models typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions whose answers provide key steps on the path toward solving the larger problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed subquestion updates beliefs about a target outcome. To automate this measurement at scale, we target as outcomes the prices in continuously updating forward-looking prediction markets. Our benchmark has three rare properties: it is contamination-resistant by construction, since ground truth is generated by future world events; it is open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and it is grounded, with informativeness measured against quantified changes in real-world beliefs (prices). We evaluate eight frontier and open-weight models on 336 target outcomes and find that VOI correlates highly with independent measures of model quality ($r=0.87$) and predicts downstream usefulness when the generated questions are passed to a separate forecasting agent.
Chat is not available.
Successful Page Load