Machine Learning for Educational Assessment of Students with Disability
Abstract
Interpretable machine learning is increasingly used to understand educational data, yet little is known about whether different machine learning algorithms produce consistent explanations of the same assessment. We investigate this question using large-scale observational assessment data from teachers of students with disability by combining Rasch modelling with Decision Trees, Random Forests, and Support Vector Machines. The study analysed two assessments targeting thinking skills and literacy, comparing students with and without autism. Machine learning models were trained to reconstruct Rasch-derived competency estimates, and feature importance rankings were compared across algorithms and groups before being interpreted alongside Rasch Differential Item Functioning (DIF) analysis. Although Random Forests achieved the highest predictive performance, the main finding was the agreement across all three machine learning models. Despite fundamentally different modelling assumptions, the algorithms consistently identified similar assessment features as the strong contributors to competency estimation. These findings closely aligned with previous Rasch DIF analyses while providing additional insight into the relative contribution of individual assessment features beyond that afforded by traditional psychometric methods. Together, the results demonstrate how interpretable machine learning can complement Rasch modelling by strengthening validity arguments and providing robust, explainable evidence about assessment structure. More broadly, the study shows that agreement across multiple interpretable machine learning models can provide independent evidence for psychometric validity while potentially generating actionable insights for educational assessment, learning analytics, and teacher decision-making for students with disability.