Understanding Antibody OOD Generalization: The Role of Mutational Coverage and Inductive Bias
Abstract
Antibody engineering campaigns proceed by mutating a reference antibody and measuring properties (e.g., binding affinity) of the resulting variants. Predictive models are used to prioritize which variants to make and measure next, so at deployment they must often score mutations unlike those seen during training. Model selection, however, is still typically performed using random in-distribution (ID) train/test splits. We show that such random ID splits can lead to misleading conclusions, and inspired by real campaigns, we propose distance-based out-of-distribution (OOD) splits for model selection instead. To characterize generalization in this setting, we introduce \emph{mutational coverage}, which quantifies the extent to which mutations in the test set, i.e., perturbations from the reference antibody appearing in the test set, are also present in the training set. Across multilayer perceptrons (MLPs), protein language models (PLMs), and graph neural networks (GNNs), reduced coverage affects performance and changes model rankings. Surprisingly, simple MLPs are very competitive at high coverage while complex PLMs underperform and GNNs perform best when coverage is limited. Our results identify mutational coverage as a key factor of model performance in distance-based OOD settings and clarify when different model inductive biases are most effective.