From Generated CIFs to Auditable DFT Endpoints: A Provenance Case Study in Li–Zr–O–Cl Crystal Generation
Abstract
Scientific AI increasingly relies on multi-stage workflows in which generated candidates are screened, simulated, and selected for downstream study. Yet final scientific artifacts often preserve what was computed while losing why a candidate was generated, screened, and selected. We present a provenance case study of a Li–Zr–O–Cl crystal-generation workflow and evaluate data readiness through task recoverability rather than artifact presence alone. From 120 generation attempts, seven within-formula top-ranked candidates were carried through variable-cell relaxation with Quantum ESPRESSO, and all reached BFGS convergence. The resulting archive preserves reproducible DFT records, including inputs, outputs, pseudopotential metadata, converged geometries, energetics, and parsed endpoints. However, two upstream provenance categories—the generator seed/checkpoint identity and per-candidate screening-decision records—remain unavailable. We analyze which scientific questions can be answered under CIF-only, archived-bundle, and complete-provenance information-access regimes, showing that endpoint verification is possible without recovering the full candidate lineage or the decisions that produced it. Our results distinguish endpoint completeness from campaign reproducibility and demonstrate that scientific data readiness should be evaluated by the tasks that available evidence can support.