Self Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
Abstract
Manually curated biomedical repositories---spanning bioactivity, genomics, and chemistry---are expensive to maintain, lag behind primary literature, and often discard experimental context. The absence of contextual information obscures critical nuances, thereby complicating the assessment of data correctness and coverage, necessary criteria for building high-quality models. We show that PubMed itself can be turned into structured datasets---autonomously and cost-effectively---that are larger, more nuanced, and more accurate than the curated databases they would replace. We present three coupled contributions: (1) an LLM-based entity-tagging pipeline, grounded in nine biomedical ontologies, that tags 4.5 billion entities across 19 categories in a 22.5M-paper, 2.5-trillion-token PubMed corpus; (2) hybrid sparse–dense retrieval infrastructure supporting surgical entity-filtered semantic queries over the tagged corpus; and (3) Starling, a multi-agent deep research system that, given only a natural-language task description, autonomously designs precision- and recall-targeted retrieval filters, induces an extraction schema, and emits structured records with nuance-rich fields and supporting passages. Applied to six tasks---blood-brain barrier permeability, oral bioavailability, acute toxicity (LD50), gene-disease associations, protein subcellular localization, and chemical reactions---Starling produces ~7.7M records (per-task scale ranges from 131K to 3M); several of these are, to our knowledge, the largest public datasets for their respective properties. Frontier-model rejection of our kept extractions is 0.6–7.7% across tasks, surprisingly far below the error rates we measure on the widely used, manually curated counterparts (e.g., 16.5% on BBBMartins, 7.3% on BioavailabilityMa). Beyond scale and accuracy, the attached supporting passages carry nuance that tabular databases discard: for example, oral bioavailability of a molecule might depend on whether the patient is fed or fasted. Together, the corpus, retrieval layer, and agent establish a foundation for multimodal predictive and generative models in AI-driven therapeutic design. We release code and links to our datasets at https://anonymous.4open.science/r/starling-B026/.