What Does a Verification Budget Buy? Sampling and Training Levers in LLM-Driven Crystal Structure Generation
Shehroz Ahmad Shoaib ⋅ Burhan SaifAddin ⋅ Trupti Mohanty ⋅ Taylor Sparks
Abstract
A generative model for crystal structure prediction does not deliver a structure; it delivers a queue of candidates, and every candidate taken seriously costs a DFT relaxation. Verification, not generation, is the scarce resource. We therefore read best-of-$N$ match rate as a \emph{verification yield}, the probability that spending $N$ relaxations on one target composition returns the right structure, and ask what raises that yield per unit of spend. We study Qwen2.5--7B-Instruct fine-tuned with LoRA to emit CIFs conditioned on reduced composition and target space group, evaluated on a frozen MPTS-52 split with \texttt{pymatgen} \emph{StructureMatcher}. Holding LoRA rank, training volume ($\sim$24k crystals) and optimizer steps (4{,}500) fixed, we change one lever at a time: a short MP-20 to MPTS-52 warm-up gives the best yield at 30.8\% against 30.0\% for direct supervised fine-tuning, a mixed corpus gives 30.4\%, and two GRPO configurations give 27.7--28.1\% while costing roughly $20\times$ more per optimizer step. The comparison that matters for verification economics is a different one: raising the budget from $N{=}1$ to $N{=}10$ moves the same baseline from 18.4\% to 30.0\%, so the sampling budget dominates every training lever we tested by an order of magnitude. We also find that using the verifier itself as an RL reward does not improve the generator once the supervised prior has converged: the continuous reward produces far richer within-group variation than the discrete one, yet held-out matching does not move. \emph{StructureMatcher} is a serviceable filter but a poor teacher at this budget.
Chat is not available.
Successful Page Load