Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?
Abstract
The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness with reliable metadata annotation. However, public repositories (PubChem BioAssay, ChEMBL) and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay-format and detection fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and flag silver label annotations directly from the assay text. We find that 42\% of records lack an assay format annotation, and over 99\% lack structured BAO annotation for the assay format and detection field. Using a PubChem-derived evaluation set, combined with labels from ChEMBL, we assess agreement with existing and LLM-generated labels. Recall exceeds 0.95 on the assay format cases, for both closed source and open-source models on the biochemical and cell-based format, but disagreements rise on under-represented classes (tissue-based, cell-free). In 81\% of cases LLMs agree with the BAO label provided by the BioAssay Research Database (BARD) instead of the third-party annotated BAO assay format. Similarly, in 39\% of disagreements between LLMs and ChEMBL annotations, these are already disagreements between ChEMBL and PubChem. We then identify cases for further review in a qualitative study involving a senior industrial curator. During discussions, LLM-generated evidence caused the expert to occasionally update their own labels, showing LLMs can flag potentially mislabeled assays. Together, these results suggest that LLMs are a useful tool for annotating assay metadata at scale, although per-class reliability estimates might be needed before such labels are used in downstream ML pipelines.