Developing Scientific Software through a Verifier-Guided Agentic Self-Improvement Loop
Abstract
The emergence of coding agents has enabled software to be developed with minimal human involvement. Verifier-guided self-improvement loops provide a framework in which agents iteratively improve software against a measurable objective. This provides an opportunity to develop specialised scientific software for tasks where performance can be evaluated automatically. We study the example of writing systematic chemical (IUPAC) names of molecules, a task for which no open-source software solution was previously available. A name can be verified as correct by using the open-source OPSIN tool to check whether it round-trips to the input structure. Using this verifier as the objective, with a development set of 2.68 million molecules, Codex GPT-5.5 implemented NISPO, an open-source Python package for IUPAC name generation, from scratch. NISPO achieves 98.1% round-trip accuracy on an evaluation set of 103 million PubChem molecules, approaching the performance of proprietary tools which have been developed over decades. We describe the framework used to develop the software, including skills for analysing failure modes and secondary objectives used to optimise output quality and code efficiency. We find that verifier-guided development is a promising approach to creating specialised scientific software; however, current agentic systems still require substantial human steering to set intermediate goals and to balance multiple objectives.