HARVEST: Opening Up the Dark Bioactivity Data of Pharmaceutical Patents with AI Agents
Abstract
Drug patents hold large tables of structure--activity data, yet these data remain computationally inaccessible: being public it's locked inside long documents that no open database has collected systematically. We present HARVEST, a pipeline of five agents that reads USPTO patent files and writes structured bioactivity records for \$0.11 per document. On 164,877 patents it produced 3.05M activity records (2.09M unique protein--ligand interactions), including 639,149 structural clusters and 1,057 protein targets absent from BindingDB, in under a week --- work that by hand would take over 50 years. Two independent checks, against BindingDB and against patents we read ourselves, both put accuracy near 0.8. Fast and reliable enough, HARVEST turns the dark bioactivity data of drug patents into open, computable data for AI-driven drug discovery.