Better Confidence, Not Better Answers: A Pre-Registered Tool-Access Ablation for LLM Variant Classification, and What It Cannot Establish
Abstract
A central open question for agentic AI in biology is when a frontier general-purpose language model needs a biology-specialized tool, and when its own parametric knowledge is already sufficient. We ran a pre-registered ablation in which the same frontier LLM (Claude, Sonnet 5) classified 90 ClinVar missense variants as Pathogenic or Benign either from parametric knowledge alone (NOTOOL) or given the same identifiers plus a real ESM2 protein-language-model score (TOOL), stratified by ClinVar review status. The pre-registered primary metric, top-1 accuracy, rose from 80.0% to 84.4%, which is inconclusive rather than null: with the observed discordance structure this design could only have reached significance at a gap of ≥11.1 points, and its simulated power against the observed +4.4-point effect is 0.10 (McNemar's exact p=0.50; bootstrap 95% CI [-5.6,+14.4] points). A post-hoc metric using the model's stated confidence as a ranking score rose from AUROC 0.863 to 0.943 (paired bootstrap 95% CI on the gap [+0.016,+0.154]), and the improvement survives restricting to the 70 items where both conditions gave the same label (0.913 to 0.971, CI [+0.013,+0.117]), so it is not merely the accuracy change in disguise. Calibration proper improves but remains poor in both arms (Brier 0.187 to 0.125; ECE 0.190 to 0.141; both arms under-confident by 13-19 points). We report three things the study cannot establish, prominently rather than defensively: (i) our two prompts were not word-matched — the NOTOOL preamble carries four clauses absent from TOOL, one of which primes uncertainty — so "tool access" is not cleanly isolated, and we bound rather than remove this confound; (ii) a zero-parameter rank fusion of the ESM2 score with the untooled prediction reaches AUROC 0.923, statistically indistinguishable from the LLM given the tool (+0.020, CI [-0.017,+0.061]), so routing the score through the LLM is not demonstrably doing work; and (iii) 82 of 90 items are the only variant sampled from their gene, so an 80% untooled baseline is compatible with recall of a public answer key. We argue the resulting picture — a specialist score that buys calibration for triage rather than accuracy — is the more useful finding, and that generalist-vs-specialist evaluations reporting only top-1 accuracy will miss it.