The TDC Lipophilicity benchmark is one assay — and a predictor that fits nothing retrieves 834 of 840 held-out labels exactly
Abstract
A predictor that fits nothing—look each compound up in ChEMBL by its ChEMBL accession, return the source-assay median—attains MAE 0.0069 on the TDC Lipophilicity benchmark's own scaffold-split test fold (TDC scaffold split, seed 42), while recent published models sit at MAE 0.42–0.54. The explanation is provenance: auditing the 4200-compound AstraZeneca lipophilicity collection against ChEMBL 36, we find that 4183 of its rows (99.60%) derive from a single assay (CHEMBL3301363), which covers 4196 molecules total—11.66 times more than the next-largest LogD assay in ChEMBL; among benchmark compounds, the runner-up assay covers only 80. All values are confined to [−1.5, 4.5], consistent with a declared instrument range. The lookup result follows from that value identity; it bounds what database retrieval achieves and describes no published model. Without the source assay, coverage falls to 26.9% of test compounds, drawn from 122 different assays, at MAE 0.764; a nearest-neighbour predictor using only the benchmark's training fold reaches MAE 0.7868, so the gap is label retrieval rather than an easy split. The benchmark is internally consistent—single laboratory, single protocol, mutually comparable labels—but evaluates prediction on one protocol, not the lipophilicity prediction task as it arises in practice.