Combining Physical and Retrieved Empirical Priors for Solid-State Synthesis Condition Prediction
Abstract
Predicting the calcination and sintering temperatures of a solid-state synthesis is an essential component of synthesis planning: while selecting precursors has received considerable attention, predicting the thermal conditions that convert them into a target has not. The task is challenging because reported synthesis data are small and span heterogeneous chemistry, leading models to struggle to generalize beyond the compositions they have seen. We investigate whether explicitly providing prior chemical knowledge can improve generalization in this data-scarce regime. We introduce CalSin-LLM, a large language model conditioned on two priors: melting point injection, which supplies target and precursor melting points as a thermodynamic-scale coordinate, and chemistry-aware retrieval, which draws structured recipes from a literature-derived corpus. Without synthetic reaction recipes, CalSin-LLM is competitive with a synthetic-data-augmented baseline on both the Kononova test set and a leakage-free 2025 ICSD test set. We further show that the two priors are complementary: neither improves both tasks in isolation, whereas their combination gives the best result in every task. Finally, because reported synthesis temperatures are quasi-discrete, we introduce tolerance-banded accuracy, with tolerance fixed a priori by the label lattice, as a decision-relevant complement to pointwise regression.