Diversity Sampling Costs Labels Near the Training Pool and Saves Them Far From It
Abstract
Choosing which molecules to label is a first-order decision whenever measurement or high-level calculation is the bottleneck. Structure-based diversity selection is the classical answer to that choice, and averaged over a test set it appears to do nothing. We show that this null conceals two effects of opposite sign. For each test molecule we compute the mean Tanimoto distance to its 100 nearest neighbours in the unlabelled pool. That quantity is available from SMILES before any label is acquired. Resolved along that distance, MaxMin selection is worse than random selection near the pool and better far from it. In all 21 combinations of extrapolation definition, model family and input representation that we test, the ratio declines with distance and falls below unity in the farthest decile. In label units the two sides are factors rather than margins. One MaxMin label is worth about 0.3 random labels in the nearest decile and about 8 in the farthest. Whether a test molecule is designated out-of-distribution does not predict the sign. Its distance to the pool does.