Resonance: From Teacher to Self-Distillation in LLMs for Reaction Prediction
Anagha Aneesh ⋅ Sebastian Ibarraran ⋅ Frank Hu ⋅ Jason Liu ⋅ Grant Rotskoff
Abstract
Building specialized models for scientific reasoning presents a fundamental tradeoff: frontier models provide strong domain capabilities, but relying on them for continued supervision limits the accessibility and adaptability of smaller open-weight models. We investigate whether frontier supervision can instead serve only as a bootstrap, using forward reaction prediction, a core task in computer-aided synthesis planning, as a testbed. \textbf{Resonance} first distills frontier-model reasoning into Qwen3-8B, then removes the teacher and continues post-training exclusively on the student's own verified and preference-ranked generations. On USPTO-50k, this raises top-1 exact-match accuracy from $1.32\%$ to $47.9\%$, statistically indistinguishable from sonnet 4.6 at $46.8\%$; notably, 17 percentage points of improvement occur after frontier supervision ends. Gains extend across reaction families and to a held-out synthetic benchmark spanning 104 reaction templates. These results show that frontier supervision can bootstrap a smaller model into a regime where it continues acquiring specialized chemical capability from its own verified experience.
Chat is not available.
Successful Page Load