When Do Compact Knowledge Graphs Suffice for Battery-Materials Hypothesis Generation?
Shashwat Sourav ⋅ Viktoriia Baibakova ⋅ Sanjay Das ⋅ Ran Elgedawy ⋅ Emily Herron ⋅ Tirthankar Ghosal
Abstract
Materials hypotheses connect a material or component, failure mode, intervention, mechanism, and measurable outcome. Knowledge graphs (KGs) make these relations explicit, but it remains unclear when compact graph context can substitute for a full graph. We study battery-materials hypothesis generation with Mistral-7B, Llama-3.1-70B, and Gemini 2.5 Flash. Beyond 100 local 15--18-triple KGs, we test candidate graphs with $N\in\{64,256,512\}$ and independently retrieved MatKG context. We define compression as successful only when top-$k$ beats size-matched random-$k$ under a paired 95\% confidence interval and retains at least 80\% of expanded-full fixed-reference recall. Among nine size-matched Top-8 settings, only Mistral-7B at $N=256$ nominally meets both criteria under an unadjusted interval; this single pass is exploratory because the nine-setting family is not multiplicity-adjusted. Independently retrieved MatKG Top-8 context shows no statistically resolved RFS advantage over Random-8 for any model. Human calibration gives $F_1=0.677$ for the lexical Relation Fidelity Score and $\rho=0.662$ between object-term KG coverage and human coverage judgments. These metrics are imperfect diagnostics rather than proof of structural reasoning or scientific quality. The resulting Conditional Compressive KG Hypothesis is falsifiable and model-, budget-, retrieval-, and task-dependent.
Chat is not available.
Successful Page Load