TabAL: Importance of Representation in Molecular Tabular Active Learning
Abstract
Molecular representation is often treated as a preprocessing choice, although it can determine whether a learner can exploit small labelled sets. We test how 1,024-bit ECFP4 fingerprints versus 210 standardised RDKit2D descriptors change pool based active learning across TabPFN v2.5, random forest, XGBoost, and LightGBM. The experiment comprises of five scaffold-split ADMET tasks, three acquisition policies, ten seeds, and 1,200 active-learning runs spanning label budgets from 32 to 256 labels. Switching to RDKit2D improves all twelve model policy macro conditions by 0.054-0.091 area under the AUROC learning curve (AULC). Switching to RDKit2D improves all twelve model-policy macro conditions by 0.054–0.091 area under the AUROC learning curve (AULC), with the largest gain observed for TabPFN under entropy-diversity acquisition. More importantly, the representation changes the comparative result: TabPFN moves from trailing the strongest tree ensembles with ECFP4 to the best aggregate condition with RDKit2D, while model rankings remain heterogeneous across individual datasets. These results show that molecular representation is not a neutral preprocessing choice in active learning; it interacts with the learner and acquisition policy strongly enough to change the conclusions of a benchmark.