More Thinking, Less Reading? Separating Reasoning Utility from Evidence Selection in Factual QA
Abstract
Reasoning modes improve many difficult multi-step tasks, but whether their additional test-time computation benefits short factual question answering or changes evidence selection remains unclear. We conduct a controlled same-model study using Qwen3-4B, comparing Think and NoThink inference under two complementary settings. First, we evaluate factual utility on 150 untouched PopQA facts that are never used for screening or selection. Thinking achieves 6.89% accuracy compared with 10.44% for NoThink, a difference of -3.56 percentage points (95% bootstrap CI [-7.78, 0.00], paired Wilcoxon p=0.078), while increasing mean generation from 6.21 to 244.53 tokens per answer—a 39.4× increase. Second, on a separate set of 150 screened-known facts, we construct aligned and counterfactual conflicting records to test whether reasoning changes reliance on explicit context versus parametric memory. Both modes achieve 100% context adherence and 0% prior reversion under the tested evidence conditions. These results separate reasoning utility from evidence selection: in this setting, substantially more reasoning computation provides no observed factual-recall benefit while explicit textual evidence remains controlling.