Simple Predictors, Strong Baselines: Revisiting Chemical Perturbation Prediction
Abstract
Predicting transcriptional responses to chemical perturbations is important for drug discovery. We revisit chemical perturbation prediction at the pseudobulk level using a lightweight MLP and benchmark complementary compound representations spanning molecular structure, induced phenotype, and mechanism of action as inputs to the predictor. Across absolute and differential metrics on the sci-Plex3 dataset over both full and DEG50 gene panel, our approach consistently outperforms control- and perturbation-mean based learning-free baselines and four representative single-cell perturbation prediction models, including ChemCPA, Biolord, Doloris, and PerturbDiff, despite using 17-55x fewer parameters. We illustrate that sophisticated molecular representations do not consistently improve prediction, whereas drug target-aware representations are particularly informative for predicting differentially expressed genes. These results highlight the potential of lightweight pseudobulk-level predictors and motivate future work exploring their use as priors for modeling single-cell response heterogeneity.