MathCD: A Benchmark Dataset for Cognitive Diagnosis with Semantic Information
Xueyi Li ⋅ Youheng Bai ⋅ Tengteng Cheng ⋅ Mingliang Hou ⋅ Teng Guo ⋅ Jiaqi Zheng ⋅ Yongdong WU ⋅ Zitao Liu
Abstract
Cognitive diagnosis (CD) aims to infer students' knowledge states from historical learning interactions, providing essential support for personalized education. Despite recent progress, most existing CD datasets mainly represent students, exercises, and knowledge concepts as discrete identifiers, offering limited semantic information about exercise text, answer choices, and student responses. This limitation restricts the development of semantic enhanced CD models and makes it difficult to evaluate whether foundation large language models (LLMs) can demonstrate educational understanding in CD scenarios. To address this gap, we introduce \emph{MathCD}, a new K-12 mathematics CD dataset collected from an online learning platform. MathCD contains three subsets with $2,197$ students, $46,431$ exercises, $3,258$ knowledge concepts, and $106,220$ student-exercise interactions. Moreover, MathCD provides rich semantic annotations, including exercise text, KC text, answer text, response text, exercise type, and exercise difficulty. Based on MathCD, we evaluate $9$ existing CD models and $5$ representative foundation LLMs. Experimental results show that semantic information consistently improves both classical and neural CD models, with semantic-enhanced variants achieving comparable or better performance than strong identifier based models. We further find that foundation LLMs exhibit emerging CD-related abilities in student response prediction and exercise difficulty discrimination, but their performance remains unstable and limited without task-specific adaptation. MathCD provides a new benchmark for semantic-enhanced CD and foundation LLMs evaluation in educational modeling.
Chat is not available.
Successful Page Load