Learner History and Calibration for Generated Mathematics Questions
Abstract
Can a language model use a pupil's practice record to predict how they will respond to a generated mathematics question? We study 1,412 paired forecasts of skill credit from 276 pupils on 439 questions in two English secondary schools. The model receives either the question context alone or the same context plus the pupil's earlier practice record. Adding history reduces Brier score (lower is better) from 0.2148 to 0.2027, a gain of 0.0121 (95% interval [0.0018, 0.0265]). This is a supervised result: each model's probabilities receive a two-parameter adjustment fitted on other pupils' recorded credit outcomes. The interval resamples pupils and questions and refits these maps. Without calibration, history still improves ranking, but its Brier gain is uncertain. Checks using separate question contexts or later dates also favour history but are inconclusive. Numerical models trained on the larger response archive are competitive. These retrospective results show that earlier responses help predict recorded skill credit. Whether this helps a tutor choose questions or improve learning needs a separate study.