SciRAC: Quality Assurance for LLM-Extracted Scientific Data
Abstract
Scientific discovery increasingly relies on data scattered across the literature to uncover scientific relationships and train predictive models. Large language models (LLMs) can extract scientific knowledge into structured datasets at scale, however, extraction errors may propagate into downstream scientific analyses. Existing methods do not always explicitly test whether the fields assembled into a record belong to a compatible experimental unit. Here, we introduce SciRAC, an agentic framework for post-extraction scientific record verification, where evidence-grounded agents repair candidate fields; an experimental-entity binding verifier assigns each populated field to its material, sample, synthesis event, and property context; and a deterministic decision layer derives a record verdict from whether these bindings form a compatible experimental instance. SciRAC improves extraction accuracy against expert annotations from 68.2% to 86.2%. Its record-level verifier flags 19.9% of the full corpus for potential misalignment; because independent record-level labels are unavailable, this is a diagnostic output rate rather than an estimate of the true defect prevalence. Across six downstream predictors of the synthesized material, supplying verified rather than raw records improve saccuracy by 7.6 to 24.7 percentage points. By making record consistency an explicit verification target, SciRAC provides a missing safeguard between literature extraction and downstream AI-for-science workflows.