Small Talk: Literature Review with Locally-Hosted Small Language Models
Abstract
Agentic literature review repeats one narrow task across a large archive, making per-document cost, privacy, and local deployment central. We ask whether lightweight, open-weight models on a laptop can rebuild a published review table from a 1,346-document corpus, and how to compare them with frontier models that may have memorized the target values. Under our threshold, claude-opus-4.8 reproduces 34 of 39 recoverable values in at least half of its closed-book attempts, whereas each local model reproduces fewer than 24\%. We control this confound with per-model closed-book probes, empirical difficulty tiers, and matched input tiers. The controls show that difficulty tiers derived from pooled arms do not transfer reliably to small models and that increasing the input budget helps some models but harms others. Nevertheless, a 20B model running both pipeline stages locally reconstructs 22 of 39 recoverable cells with no frontier model in the chain. Three of four local models are bit-deterministic at temperature 0 in repeated closed-book probes, whereas gateway-served arms are not; served-arm error bars therefore include variability from the serving stack. These results support compact models for the repetitive bulk of literature review, provided input budgets are selected per model and provenance, rather than raw recall, governs deployment.