Random Splits Overstate Translational Performance Across Nanoparticle Formulation Benchmarks
Zhongjun Zheng ⋅ Zhirun Yue
Abstract
Machine learning could reduce experimental screening in nanoparticle formulation, but its reported performance is usually measured with a Random split. Literature-curated benchmarks are hierarchical: formulations from the same molecule and paper share chemistry, protocols, equipment, and other unrecorded conditions. A Random split can therefore place related observations in both training and test sets, measuring interpolation without establishing transfer to a new payload, study, or chemical neighborhood. We test this distinction in two public benchmarks comprising 433 poly(lactide-co-glycolide) formulations and 1,092 lipid nanoparticle records. Alongside the Random split, we evaluate three deployment-oriented protocols: a Molecule- or Ionizable-component-grouped split, a Study-grouped split, and, where structures are sufficiently available, a Structure-clustered split. Random forests and XGBoost predict particle size and encapsulation efficiency under ten repeated five-fold evaluations; median, ridge, and k-nearest-neighbor baselines, temporal holdouts, duplicate controls, representation ablations, conformal intervals, and label-provenance checks test robustness. Random-forest $R^2$ for PLGA size/encapsulation efficiency falls from 0.713/0.598 under the Random split to -0.030/0.130 when molecules are held out. For lipid nanoparticles, 0.464/0.450 falls to -0.029/0.114 for unseen ionizable components and 0.001/-0.011 across studies. Future-publication size $R^2$ remains negative at three temporal cutoffs; deduplication and SMILES descriptors do not remove the transfer gap, while nominal conformal coverage requires impractically broad intervals. These results support a methodological conclusion rather than a new predictor: Random splits remain informative for within-platform interpolation, but grouped, temporal, and provenance-aware evaluation is necessary before claiming translational generalization.
Chat is not available.
Successful Page Load