DeepPURPLE: Open-Vocabulary Peptide Design in Non-Canonical Amino Acid Chemical Space
Abstract
Non-canonical amino acids (NCAAs) are widely used building blocks in peptide therapeutics, introducing chemical functionality beyond that of the 20 canonical amino acids. However, experimentally resolved NCAA-containing peptide–protein complexes are scarce, limiting the data available to learn NCAA-specific structural compatibility at interfaces. Due to this limitation, most sequence design models are restricted to pre-defined amino acid types, leaving unseen amino acid types outside the design space. We present DeepPURPLE, an NCAA-compatible peptide sequence design framework, trained on a synthetic dataset generated by Rosetta sampling integrated with CHARMM36m NCAA force field, to predict a residue-wise embedding for generic NCAAs with side-chain modifications. We constructed the synthetic dataset spanning 341 side-chain-modified NCAA types through systematic NCAA substitution, structure optimization, and physicochemical filtering. This open-vocabulary peptide design formulation enables zero-shot NCAA design without extra training by retrieving amino acids close to the predicted embedding in a pretrained chemical embedding space. DeepPURPLE was assessed on a held-out crystal test set for its transferability to real data and to unseen NCAAs, and achieved top-20 recovery of 20.4% at seen NCAAs, 9.7% at unseen NCAAs and 43.9% at canonical amino acids. Overall, this work provides a scalable computational framework for peptide sequence design in a chemically diverse amino acid space.