Discrete Diffusion for Protein-Conditioned Coding-Sequence Design and Reward Alignment
Abstract
A single protein can be encoded by many synonymous mRNA coding sequences (CDSs), whose sequence composition affects RNA structure, stability, and protein output. We introduce Prot2RNA, a protein-conditioned discrete diffusion language model (DLM) for human CDS generation. Under matched architecture, data, and training budget at this model scale, diffusion finetuning outperforms next-token prediction (NTP) and masked language modeling (MLM) alternatives, while a modality-aligned positional encoding further improves conditional codon modeling. We then apply reinforcement finetuning to Prot2RNA using Group Relative Policy Optimization with normalized minimum free energy as a sequence-level objective. Reinforcement finetuning consistently increases predicted structural stability, with KL regularization modulating policy adaptation and sequence diversity. Finally, experimental evaluation in a cell-based reporter system shows higher reporter output for the tested reward-finetuned designs than for the tested unaligned Prot2RNA design. These results support discrete diffusion as a framework for learning a protein-conditioned CDS prior and subsequently adapting it toward application-specific sequence properties.