Antibody Optimization via Reinforcement Learning over a Learned Somatic Hypermutation Prior
Abstract
Antibody affinity maturation is a natural optimization loop: B cells propose mutations through somatic hypermutation (SHM), and variants are selected by antigen-binding affinity. Modeling both halves jointly is difficult, because mutation models capture SHM sequence patterns without optimizing binding, while affinity optimization can improve binding along trajectories that are not naturally accessible. To couple these two halves, we introduce \textit{iGC} (\textit{in silico} Germinal Centers), a two-stage reinforcement learning (RL) framework. Stage 1 learns an SHM mutation prior from parent--child pairs in B-cell lineage trees and domain-adapts it to the target affinity-maturation dataset. Stage 2 initializes an RL policy from this prior and optimizes antigen binding while remaining KL-regularized toward a frozen copy of it. Because raw affinity rewards drive the policy onto a few dominant substitutions or structurally tolerated framework sites, we introduce a Reward--Tax Meta-Gradient (RTMG) objective: RewardNet smooths the affinity landscape, TaxNet discourages overexploitation, and their difference is a learned meta-reward. The Stage~1 prior more than doubles the strongest baseline on mutation-site prediction, and after RL optimization of iGC-predicted mutation positions still lie in WRC hotspots against a random background. iGC attains the highest reward--mutation-diversity product among all compared methods. Code and implementation details are available at \url{https://anonymous.4open.science/r/iGC-0A08/README.md}.