K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Abstract
Large language models (LLMs) are increasingly deployed in K–12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval measure only whether a model can answer an exam question, i.e., factual recall. Effective educational AI further requires curriculum cognition: the structured understanding of how knowledge is organized, including prerequisite chains, concept taxonomies, experiment–concept links, and pedagogical sequencing. Curriculum cognition is neither probed by current benchmarks nor explicitly taught by current instruction-tuning data. To close this gap, we introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from the official People's Education Press textbooks, covering mathematics, physics, chemistry, and biology across primary, middle, and high school, with seven node types (Concept, Skill, Experiment, Exercise, Section, Chapter, Book) and nine relation types spanning taxonomy, prerequisite, association, verification, assessment, location, and order. Building on this single graph, we derive two complementary resources: K12-Bench, a 23,640-question multi-select benchmark across five graph-derived task families (Ground, Prereq, Neighbor, Evidence, and Locate) that jointly probe curriculum cognition; and K12-Train, a KG-guided supervised fine-tuning corpus of ~2,300 QA pairs synthesized from node attributes and edge semantics. Experiments expose a clear gap and a clear remedy: on K12-Bench, even the strongest proprietary model (Gemini-3-Flash) reaches only 57% exact match and the best open-source model (Gemma-4-31B-IT) only 46%, with Prereq and Neighbor being the hardest; yet under a strictly matched 2,300-sample SFT budget on Qwen3-4B-Base and Llama3.1-8B-Base, K12-Train consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora (OpenHermes, Infinity, UltraChat, WizardLM, DataFlow, LMSYS, SmolTalk, Tulu-3) on both GaokaoBench and EduEval, showing that structural curriculum grounding is remarkably sample-efficient for educational SFT. We release the graph, benchmark, training data, and the full construction pipeline.