Obscure or Invented? Auditing Gene Identifier Knowledge Across Twelve Language Models
Abstract
Language model agents are beginning to run applied gene discovery workflows, where they answer questions about gene identifiers and decide when to verify a claim against a database. Models internally represent whether they recognize an entity, but prior work covers celebrities and cities rather than the dense biomedical vocabularies agents act on, where nearly every plausible string could name a real gene. We probe ten instruction-tuned open-weight models (0.5B to 27B parameters) on a citation-matched benchmark of well-studied, understudied, and fabricated human gene symbols. Every model carries a graded linear signal of how well studied a real gene is, the signal sharpens with scale, and no model separates fabricated symbols from real understudied ones at any scale. The familiarity signal is most readable exactly where a prompt asks the model to answer or declare non-recognition, and the decision barely uses it. We also document an evaluation trap, HGNC nomenclature records study status, so a probe evaluated without a surface-form baseline can appear near perfect while measuring orthography. Verification in gene discovery agents should be enforced by the harness or read from activations, never delegated to the model's own sense of recognition.