Does the Proxy Preserve the Decision? A Stopping-Policy Audit of Multiple-Choice Reformulations of Incremental QA
Abstract
Many AI benchmarks make evaluation cheaper by replacing an open-ended task with a fixed set of answers. This paper asks whether that change can also alter when a system appears ready to act. Quizbowl provides a clean test case: questions arrive clue by clue, so a model must decide not only what to answer but when to stop waiting for more evidence. We compare open-ended and four-choice versions of the same 3,037 questions across 96 prespecified stopping-policy specifications. The primary result is simple: for the median question, the stopping point does not move, and the family bootstrap interval is [0,0]. But this does not establish that the two formats are equivalent. Depending on how probabilities are calibrated and how future clues are valued, between 52% and almost 100% of questions keep the same stopping point, while average shifts range from about 0.68 clues earlier under multiple choice to 0.11 clues later. Under isotonic calibration, most specifications favor earlier stopping with multiple choice. The broader lesson is that a constrained answer set can leave the typical question unchanged while still moving a consequential minority. That distinction matters anywhere a model must decide whether to answer, alert, escalate, or wait for more information: final accuracy, or even a zero median timing shift, is not enough to show that a simplified benchmark preserves the same decision policy.