SprayBench: Benchmarking LLMs on Open-Ended Materials Reasoning in Additive Manufacturing
Abstract
While LLMs are increasingly used in scientific workflows, limited evidence exists to date that they can support materials researchers and process engineers in open-ended manufacturing decisions that require reasoning over material systems, processing conditions, process-property tradeoffs, and sparse experimental evidence. We introduce SprayBench, a 525-question benchmark for open-ended reasoning in the emerging field of cold spray materials. SprayBench combines source-anchored literature extraction with cold-spray expert refinement to create practitioner-style questions targeting explanation, tradeoff analysis, and process guidance. Across 34 frontier and open-weight configurations, with open-weight models run locally on a DGX Spark to reflect lab-scale deployment, we explore three research questions: how well models answer without retrieval, how much source-controlled RAG helps, and how reliable LLM-as-a-judge protocols are for scalable open-ended grading. We find that, under closed-book inference, open-weight pass rates range from 6.3% to 65.8%, while the strongest frontier model reaches 85.1%. Source-excluded RAG improves all tested configurations by +3.1 to +29.4 percentage points, even when the originating document is withheld. Repeated LLM judging and a 105-response stratified expert audit quantify reviewer noise, model drift, and pass-fail threshold effects in open-ended grading.