MCP-CLIP: Composing Multi-View Representations for Cyclic-Peptide Context Retrieval
Abstract
Cyclic peptides are promising drug modalities, and generative models can produce candidates at a scale that makes exhaustive docking or co-folding impractical. FASTA, HELM, SMILES, molecular-graph, and 3D-conformer views describe the same peptide in different ways, but combining all of them may introduce redundancy or interference. We introduce MCP-CLIP (multi-view cyclic peptide CLIP), which aligns protein pockets and peptides in a shared embedding space, and use 2.66 million synthetic CPSea pairs to compare compositions of five frozen peptide view–encoder branches under a matched Stage-2 protocol. Evaluation is source-context recovery on matched CPSea pairs, a development proxy rather than binding or prospective screening. In a 54,266-pair development gallery, adding HELM to the 3D peptide representation raises target-to-peptide R@10 from 0.8358 to 0.8478 under the same target tower and protocol. On a post-hoc 4,413-query slice excluding direct Stage-1 selection and known training-shared provenance and peptide identity, the gain is +0.0055 on average and positive in every seed, although its 95% cluster-bootstrap interval includes zero. All five views together are near-tied with 3D on the full gallery, at R@10 0.8362, and trail it by 0.0394 on that slice, with an interval entirely below zero. On PAMPA membrane permeability, a target-free endpoint, CPSea-adapted SMILES+graph reaches mean Spearman ρ = 0.6885 against 0.6664 from public initialization. Together, these results favor deliberate view composition over indiscriminate aggregation under the tested protocol, with modest transfer of CPSea adaptation to a measured property.