Building on Data-Efficient QSAR:A Modular Residual Adapter with Pretrained Molecular Views for Low-Data ADME Prediction
Abstract
Reliable ADME prediction is important for compound prioritization in drug discovery, yet many endpoints are supported by limited labeled data, making data-efficient learning critical for reliable prediction. Feature-based QSAR models built on molecular fingerprints and physicochemical descriptors remain strong in this regime, while pretrained molecular models provide complementary representations learned from large-scale external data. Integrating these two sources of information is therefore a natural strategy for improving low-data ADME prediction. However, under limited endpoint supervision, directly learning how to combine multiple representations can be unreliable, causing fusion gains to vary substantially across endpoints and model designs.We address this problem with a modular residual adapter that retains a strong QSAR model as the predictive base, uses multiple frozen pretrained molecular views to learn its remaining errors, and combines their incremental corrections through constrained fusion.Because the adapter interacts with the base only through its predictions, it can be attached to different QSAR models without modifying their internal architectures. When reliable incremental evidence is absent, the adapter retains the base prediction. Across eight ADME endpoints, the adapter improves the mean performance of an XGBoost base on every endpoint and achieves the best mean performance among the evaluated methods on six. We further identify two complementary gain patterns: correction dominated by a single strong view and joint improvement from multiple individually weak views. Consistent with this plug-and-play design, the same framework produces improvements across multiple endpoints with both XGBoost and MapLight bases while shrinking its correction when little residual signal remains. These results show that constrained residual adaptation can reliably and effectively use pretrained molecular representations to enhance low-data ADME prediction.