Curriculum, Corpus, or RL? Matched-Budget Training Levers for LLM-Driven Crystal Structure Generation with Qwen
Shehroz Ahmad Shoaib ⋅ Burhan SaifAddin ⋅ Trupti Mohanty ⋅ Taylor Sparks
Abstract
Large language models can be fine-tuned to generate crystallographic information files (CIFs), but it remains unclear whether accuracy is best bought through data ordering, corpus composition, reinforcement learning, or greater adapter capacity. We study Qwen2.5--7B-Instruct, fine-tuned with LoRA to generate CIFs conditioned on reduced composition and target space group, and evaluate on a fixed held-out MPTS-52 test set using \texttt{pymatgen} \emph{StructureMatcher}. Holding LoRA rank, training volume ($\sim$24k crystals) and optimizer-step budget (4{,}500 steps) fixed, we change exactly one lever at a time. A short MP-20 to MPTS-52 warm-up gives the highest best-of-10 match rate among the fixed-capacity strategies, improving on direct MPTS-52 supervised fine-tuning from 30.0\% to 30.8\%; a mixed MP-20/MPTS-52 corpus reaches 30.4\%; and two GRPO configurations reach 27.7--28.1\%, underperforming supervised fine-tuning while costing roughly $20\times$ more per optimizer step. Logging every rollout shows why: the continuous reward does produce far richer within-group reward variation than the discrete one, yet held-out matching does not move, so the bottleneck is the match objective rather than the shape of the reward. In a separate, non-comparable sweep, raising the LoRA rank improves accuracy monotonically through $r{=}128$ with no sign of saturation. The effects between fixed-capacity levers are small (1--2 percentage points) and we report them as directional evidence about \emph{where} the next percentage point comes from under a fixed budget, not as a leaderboard claim.
Chat is not available.
Successful Page Load