Beyond Schema Conformance with Executable Verification for Scientific Configuration Repair
Abstract
Large language models can generate structurally valid scientific configuration files, but whether those configurations reproduce the intended computation is rarely tested. We study this gap on Palace, a finite-element solver for computational electromagnetics, using an executable verification hierarchy that measures schema conformance, operational acceptance by the solver, and numerical reference conformance. We construct a source-grounded fine-tuning corpus and a configuration-repair benchmark containing both shipped cases represented in the training corpus and newly authored cases created after the model and retrieval index were frozen. Fine-tuning provides the largest improvement, while the RAG-proxy setting and schema-guided repair provide additional gains. However, these improvements transfer much more strongly at the structural level than at the numerical level. On newly authored cases, the full pipeline reaches 89\% schema conformance and 81\% solver acceptance, but only 41\% reference conformance. We further show that some failures arise from insufficient task specification and that blind LLM judges do not consistently reproduce execution-based labels. These results suggest that scientific configuration generation and repair should be evaluated against executable numerical references rather than schema conformance or model judgment alone.