How Post-Training Shapes Biological Reasoning Models
Abstract
Post-training is central to building scientific reasoning models, but its effects remain poorly understood in biology, where models integrate natural language with multimodal biological data. We study how post-training stages shape generalization, and when they improve performance versus induce over-specialization. Across genomics, transcriptomics, and proteomics, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that post-training induces stage-specific trade-offs rather than uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that scientific reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed epoch-level budget, the best ID-OOD trade-off comes from combining minimal SFT with larger RL budgets and allocating adaptation capacity asymmetrically across stages.