Beyond Number Extraction: Scientific Fidelity of Local Language Models for XRD Crystallite-Size Extraction
Abstract
Materials-science literature contains characterization values whose scientific meaning depends on sample identity, measurement technique, analysis method, and provenance. Extracting the correct number while losing these relations can silently corrupt downstream datasets. We present a five-paper diagnostic study of locally runnable Qwen3.5 models (2B, 4B, and 9B) for extracting explicitly reported X-ray diffraction (XRD)-derived crystallite sizes. Thirty-six gold records were constructed through independent candidate extraction followed by manual adjudication against the source papers. Under a fixed prompt, seed, and faithful full-text condition, core selection improved with model scale: 2B achieved 90.9% precision and 55.6% recall, 4B achieved 72.0% precision and 100% recall, and 9B achieved 100% precision and recall on the pilot. The error profile, however, changed qualitatively rather than disappearing monotonically. The 2B model showed omissions, relation-binding errors, and repetition; 4B recovered every target record but leaked non-XRD measurements on two papers; and 9B retained perfect core selection while still losing method-attribution and uncertainty information. A table-serialization control further showed that document representation can create apparent model errors. These results motivate materials-specific evaluation that separates numerical retrieval from relation, scope, provenance, and representation fidelity.