Knowing What to Trust: Calibrated Evidence Scores for an LLM-Constructed Fuel Knowledge Graph
Yanpeng Ye ⋅ Qi Wang ⋅ Mani Sarathy
Abstract
Agentic systems for molecular science ground themselves in chemical databases, but a database built by a language model inherits the extractor's failure modes. We ask whether that unreliability can be made predictable rather than merely reduced. We construct a fuel knowledge graph (FKG) from about 40,000 combustion articles using an open-source extraction pipeline, and validate it against the manually curated ReSpecTh database over 277 overlapping papers, recovering 0.75 to 0.87 of the curated content while covering nearly two orders of magnitude more papers and fuels. Because completion runs over graph neighbourhoods rather than a fitted global model, every completed record carries an evidence score computed from the agreement among the neighbours that support it. On ignition delay time completion this score is calibrated: error decreases monotonically across evidence intervals, the rank correlation between evidence and error is -0.44, and test R$^2$ rises from 0.850 overall to 0.964 on the high-confidence subset. Tracing which records support each prediction shows the graph recovers the reference-fuel hierarchy of the combustion literature without being told it exists. The graph therefore reports not only retrieved values but where those values can be trusted.
Chat is not available.
Successful Page Load