Do LLMs Have Research Taste? An Analysis Using BóLèBench
Abstract
LLMs are increasingly used for recursive self-improvement research, however, compute is finite, so the harder question comes first: which idea to pursue, decided before any result exists. For human researchers this is a matter of taste, built from years of watching ideas succeed and fail. When LLM judges make this call and allocate real compute, that taste has never been measured directly: existing evaluations grade it against publication records, which hide failures and reward memorization. We introduce BóLèBench, a benchmark built on a simple inversion: use hindsight to grade foresight. We start from an existing benchmark in which coding agents built on different LLMs attempted the same ten ML research tasks, so every task already has many finished, scored attempts. A judge reads two anonymized plans for the same task, with all results removed, and predicts which attempt scored higher. The executed results are the answer key, so no human annotation is needed. Our contribution is the testbed itself: rules for which corpora and which pairs make fair questions, five kinds of research decisions to grade, and a leaderboard that scores any judge, model or human. Items are rebuilt from fresh, post-cutoff corpora as models retrain. The findings are sobering: (1) a one-line pick-the-bolder heuristic beats every frontier judge we test; (2) every judge falls below chance on dark-horse winners; (3) the judges miss the same pairs far too often to count as independent opinions; and (4) used to select one attempt from many, the weakest judge does worse than random choice.