The Anchoring Property: What One Clean Biomedical Catalogue Reveals About Provenance in Three Other Domains
Abstract
The NCBI Pathogen Detection Reference Gene Catalog passes a provenance audit that three other scientific compilations fail, and the reason is structural rather than editorial. Its records are anchored to sequence accessions, so 56% carry no literature citation at all, while every one of the 2,827 citations it does carry resolves. We ask whether that anchoring property holds outside biomedicine, and find that it does not. Applying a single instrument—does the stated source resolve, does it support the claim, are the values needed to reuse the record present—with Q1 and Q3 as a census across 32,705 records in four unrelated domains, we find that 68.2% of 18,381 defence technical reports cross-reference a catalogue that is not resolvable from the public internet; that 26.3% of 1,557 critical-mineral records have provenance that is absent or unobtainable, in an inventory whose median newest cited year is 1977; and that a crop-trait ontology citing references for 66.8% of its variables reaches a resolvable identifier for only 12.2% of those, or 8.2% of all variables. We state the implied upper bound on what a citation-following agent can verify in each corpus, and show that a readiness metric scoring citation presence alone ranks the cleanest corpus in the study as the worst.