When More Domain Data Is Not Enough: Continual Pretraining for Protected-Peptide Solubility
Abstract
Liquid-phase peptide synthesis (LPPS) offers a scalable route to peptide manufacturing, but each protected intermediate must remain soluble during reaction and recoverable during purification. The protected-peptide chemistries that govern this materials decision are sparse in standard pretraining corpora, and it remains unclear whether scaling domain-specific continual pretraining (CPT) improves downstream scientific utility. We construct one- and three-million-molecule protected-peptide corpora, expand PeptideCLM-2's molecular-property objective from 99 to 104 targets, train ten CPT checkpoints with replay, and evaluate the three-million-molecule models on PeptSol, an open benchmark of 1,388 peptide--solvent measurements. Exact same-start comparisons are architecture dependent: ChemBERTa has small positive estimates, whereas native-hybrid PeptideCLM-2 becomes negative after patent augmentation; because corpus-specific validation sets differ and smaller checkpoints lack downstream evaluation, the experiments do not establish a scaling law. These results show that lower pretraining loss alone is insufficient evidence of better materials reasoning and that scaling studies must hold validation distributions fixed and measure downstream utility at every scale.