Language Model Verifiers Fail Where Chemical Entity Attribution Errors Concentrate
Abstract
AI scientists could accelerate drug discovery and chemical research by reasoning over external knowledge. Yet this promise rests on a basic, under-tested capability: binding retrieved knowledge to the correct entity. We refer to this capability as entity attribution and study its failures in chemistry, where compounds can be difficult to distinguish because of both structural relationships and naming conventions. To separate their contributions to attribution errors, we construct a controlled stress test that varies these two sources of confusion independently. Across 526 compounds and nine open-weight language models, attribution error rates range from 37.2\% to 69.8\%. Errors are highly structured: distractors close in both structure and name account for 40.9–57.6\% of failures, despite representing only 25\% of distractors. We then ask whether verification can catch these mistakes. We convert attribution errors into false claims about the target compound, prompt language models to verify each claim, and test if they can reject those false claims. Among the eight models that regularly reject false claims, rejection rates for claims derived from the most confusable errors range from 6.1\% to 46.7\%, with even the best unable to reject the majority. Thus, the errors models are most likely to make are also among the hardest for them to detect. Examining attribution errors and verification together suggests that evaluations of verifier based on arbitrary false claims may overestimate their reliability in AI-scientist workflows.