I Can't Believe Domain Pretraining Is Not Better: Inconsistent Gains for Protected-Peptide Solubility
Abstract
Continual pretraining (CPT) is a plausible remedy when a biological foundation model is repurposed for therapeutic peptide development but has seen little of the relevant process chemistry. For protected peptide intermediates, however, molecule-only pretraining may not capture the peptide--solvent interaction needed for synthesis and isolation, and common scaffold splits may overstate generalization to new peptide candidates. We generate one- and three-million-molecule protected-peptide corpora, expand PeptideCLM-2's molecular regression objective from 99 to 104 targets, train ten CPT checkpoints with replay, and evaluate the three-million-molecule models as frozen encoders on 1,388 peptide--solvent measurements. Exact same-start comparisons are architecture dependent: ChemBERTa has small positive estimates, whereas native-hybrid PeptideCLM-2 becomes negative after patent augmentation; a positive protic-solvent slice under scaffold splitting also reverses under peptide-held-out splitting. The absence of a consistent gain exposes objective and evaluation mismatch as plausible, actionable failure modes, motivating interaction-aware pretraining, fixed validation distributions, and downstream tests at every scale.