PriorBPB: Reducing the Answer Length Confound of Bits-per-Byte While Improving Accuracy Prediction
Tom Burns ⋅ Prabhu Teja Sivaprasad ⋅ Frank Schneider ⋅ Martin Reinhardt
Abstract
Bits-per-byte (BPB) normalises log-likelihood (in bits) by bytes rather than tokens to stay tokeniser-independent, and is widely used to compare language models across heterogeneous tasks. It also depends on how long the answer strings are. The bits a model spends on a gold string can be written as $\beta + \alpha/n + \varepsilon$, with $\alpha$ a fixed overhead, $\beta$ a per-byte rate, $(\alpha,\beta)$ the least-squares fit in length, and $\varepsilon$ the rest. Since the residuals sum to zero, total bits over total bytes is exactly $\beta + \alpha/\bar{n}$, with $\bar{n}$ the task's mean gold length. Because costs that do not scale with length are spread over the answer, tasks with long answers score better, as do long items within a task. Averaging per-item BPB has the same dependence. Across a 19-task suite on Qwen3 models, BPB correlates with mean gold length at Spearman $-0.82$ to $-0.86$. OLMo-3-7B behaves the same way (Spearman $-0.79$). Among the Qwen3 models we tested, the smallest model appears to score best (an apparent inverse scaling behaviour) due primarily to differences in the overhead. We propose a less length-biased metric, PriorBPB, which estimates a per-(model,task) marginal per-byte cost curve from per-token surprisal and averages it under one shared length prior instead of under each task's own length distribution. This cuts the mean absolute Spearman correlation with answer length from $0.68$ to $0.18$ across 60 base checkpoints from five model families, while correlating with accuracy as well or better on base and instruct checkpoints (mean Spearman $0.50 \rightarrow 0.54$).
Chat is not available.
Successful Page Load