Ensemble Performance at Single-Model Cost: Offline Pseudo-Label Distillation for ADMET Prediction
Davide Boldini ⋅ Lukas Friedrich ⋅ Daniel Kuhn
Abstract
Foundation-model ensembles are strong ADMET predictors but require multiple model evaluations at inference. We distil a diverse $7$-teacher ensemble into single students by adding sparse task-level pseudo-labels, selected from $3.2$M candidate compounds, to multi-task post-training before benchmark fine-tuning. We evaluate four selection strategies and four pool sizes across seven student architectures on the ExpansionRx and PXR openADMET benchmarks. Under test-set selection among $16$ configurations per student, the best distilled models achieve RAE $0.564$ on ExpansionRx (ensemble: $0.602$; leaderboard rank $31\!\to\!7$) and $0.569$ on PXR. Best augmented students match or undercut the ensemble RAE for $5/7$ architectures on ExpansionRx and $6/7$ on PXR, using one model invocation per task group instead of seven. Architecture-level timing proxies estimate a $2.1\times$ serial speedup for GPS3D. No single selection strategy dominates. Our student-specific \emph{delta} strategy, which ranks compounds by student--ensemble disagreement, yields its largest gains on the most out-of-distribution compounds (reducing RAE by at least $0.07$ in the hardest Tanimoto decile $Q_1$ for $9/14$ architecture-benchmark pairs). These results establish offline pseudo-label distillation as a lower-cost route to ADMET prediction.
Chat is not available.
Successful Page Load