An LLM-Curated Self-Improvement Loop for ADMET Models, and the Leakage Census Its Evaluation Requires
Abstract
We implement the loop this workshop’s Track I describes — a general-purpose language model iteratively curating public training data for a bio-specialized model under automated benchmark feedback — and find the hard part is not building it but scoring it. Our loop runs 3 rounds × 3 seeds × (two local curators and a random-curator control) over ChEMBL 36 for a solubility predictor, with a scaffold hold-out committed by SHA-256 before the first call. Under a pre-registration fixed before any data existed, we first built the instrument that makes its score readable: an exact nearest-neighbour census of 46,276,308,000 Tanimoto comparisons against all 2,854,800 ChEMBL 36 structures. 93.84%, 99.52% and 50.78% of the scaffold-split test molecules of BBB-Martins, Lipophilicity-AstraZeneca and Solubility-AqSolDB have a near-duplicate (Tanimoto at least 0.9) in the curator’s own corpus, against 0.25%, 2.14% and 1.50% in each benchmark’s own training split. On the one benchmark with a usable leakage-clean subset, a ChEMBL-trained forest scores 1.5432 MAE on near-duplicated test molecules against 1.9524 on the rest (+0.4093 [+0.2949, +0.5262]). Our pre-declared falsifier did not fire — a near-duplicate carries a logD label 0.0808 log units from the benchmark’s own, against 1.6181 for random pairs — but matching cuts the gap to +0.1595 and a paired ChEMBL-specific residual of +0.0811 [+0.0104, +0.1567]: most of the raw effect is composition, against the strongest form of our own claim. Scored against it, the loop’s reported improvement exceeds its leakage-clean improvement by +0.0216 MAE [+0.0085, +0.0353], against +0.0021 for the random control — yet every arm loses accuracy on the pre-committed hold-out. Two pre-registered audits locate the bias: the curators do not steer toward reproduced chemistry, and shuffling the feedback column changes which clusters they pick (Jaccard 0.2494) but not where the loop ends (+0.0067 on the hold-out, spanning zero). The overstatement is a property of the evaluation split, not of curator behaviour, and not inflated gains. We close with six design rules for building and scoring such loops.