Mind the Tokenizer: Why a Single Space Decides What Multiple-Choice Evaluations Measure
Abstract
A common convention in multiple-choice evaluations of pre-trained LLMs is to end the prompt with "Answer:" and score the model on the likelihood it assigns to each option label (e.g., "A", "B"). As in ordinary text, a space separates the colon and the label, and it has to be part of either the prompt or the answer, yielding two possible splits of the same string. Sanz-Guerrero et al. find that this seemingly cosmetic choice changes accuracy and even reorders model rankings, and recommend the consistently better-scoring split as the default. Why the space matters at all and whether the better-scoring split is the more meaningful one has not been explained. We trace both to a hidden assumption about the composability of the tokenizer. Common tokenizers violate this composability exactly at the space, and only the better-scoring split reproduces the tokenization of the whole string, i.e., the one the model saw during training. Empirically, violating composability reduces the likelihood for of the correct answer by several orders of magnitude for three recent base models, while leaving accuracy largely unchanged. We argue the choice of the split in benchmarking should be informed by whether the intention is to test a model's knowledge or its robustness to minor input variations.