Forced Choice Is Not a Tolerance What Multiple-Choice Grading Accepts on BixBench's Numeric Questions
Vijayavallabh J
Abstract
BixBench grades an agent's free-text answer open-ended and also through multiple choice: a second model maps the answer to one of four options, forced to choose or allowed to refuse, and the forced-choice score is still reported in 2026. On the benchmark's numeric questions we set both gradings beside a tolerance --- is the submitted number within $5\%$ of the key? --- for its published runs, for open-weight agents and for three current agents. Forced to choose, BixBench's grader accepts a fifth to a quarter of the wrong answers, about what a random choice among four would, but not at random: it accepts the misses whose nearest option is the key several times as often as those nearest a distractor. The classical correction for guessing, which assumes a random choice, therefore agrees with the tolerance on the published runs only because these errors cancel, and falls below it on the current release; on current agents forced choice still adds $11$ to $16$ points. The options also leak the key: v1.5 places it second-smallest on $51\%$ of numeric items, a rank that a rule ignoring the question learns on held-out capsules. For a grader that follows the nearest option, hiding the rank and rejecting misses conflict, because a uniformly drawn rank leaves keys at an extreme, where every miss beyond the key is nearest it; the cost is large for graders shown only the answer and the options and small for BixBench's, which read the notebook. Numeric answers should be graded by tolerance.
Chat is not available.
Successful Page Load