How Fine Is the Grid? A Resolvability Law for Molecular Property Benchmarks
Abstract
MoleculeNet leaderboards routinely rank molecular representation learners by ROC-AUC margins below 0.01. We ask the prior question: can the benchmark resolve the difference? We combine Hanley--McNeil AUC variance with the DeLong pairing term to express a paired minimum detectable effect (MDE) in test size, class balance, and between-model correlation. The law is calibrated on 240 trained-model pairs, checked by per-benchmark bootstrap and seed variance, and measured across five decades of test size. Across the plausible correlation range, no published top-two gap on six MoleculeNet classification tasks is resolvable at r <= 0.9; only ClinTox crosses at r approximately 0.94. A selection analysis is sharper: on four tasks the top-five ordering is indistinguishable from a permutation of a tie, while the probability that the reported leader is truly best is 0.27 on HIV and 0.35 on BACE. A training-free composition of fingerprints and frozen molecular-language-model read-outs falls inside the best incumbent's 95% interval on five tasks, as a check rather than a claimed win. The inverse law says that today's sub-0.002 margins require roughly 10^5--10^7 test molecules. The contribution is an admissibility-and-selection screen and its verdict, not a new variance estimator: current MoleculeNet classification rankings are statistically indistinguishable from noise.