Towards Characterizing Question Format Bottlenecks for Fine-Tuned Small Language Models
Abstract
The same asymmetry appears on several public benchmarks: four-way MC is learned on SciQ and MedMCQA, while binary QA shows significantly lower improvements from chance on Mintaka and StrategyQA throughout our model range. While this phenomenon—that format affects measured capability—is known at larger scale, we show that this effect survives format-specific fine-tuning, persists to 3B, and is largely removed by context rather than by scale. Supplying the relevant context passage fully recovers binary accuracy to 0.79–0.89 on WikiDoc, whereas removing it or replacing it with an unrelated passage returns performance close to closed-book, indicating that the verification task is not unfeasible for the model when the evidence is available in context. We ask what could explain this behavior and report several preliminary diagnostics that rule out simple explanations. A global yes/no threshold correction does not repair the failure, and a preliminary last-token linear probe exposes no more answer signal than the model's own output margin. Although non-trivial data confounders cannot be excluded, in the public benchmarks as much as in ours, whether the remaining bottleneck lies in acquisition, retrieval, or the decision learned by SFT is left open.