Benchmarking safety and pedagogy in AI tutoring tools: turning policy into dual-track evaluation
Abstract
When children use AI tutoring tools, two categories of risk arise: heterogeneous safety harms and chronic pedagogical harms. We describe a dual-track benchmark that considers both categories and their interaction. The benchmark operationalises policy constructs set by the UK Department for Education's (DfE) published guidance into testable measures and uses complementary human and synthetic evaluation strands that respect the ethical constraints imposed by the two risk categories. Benchmark scores are defined as fitted values from hierarchical GLMs based on nested, repeated-measures data rather than traditional weighted aggregates. We present the design of our benchmark (withholding some details to prevent gaming) and the policy context. As annotation is currently in progress, scores cannot yet be reported. Therefore, we instead focus on the open discussion points arising from our work including the translation of policy into measures, best practices concerning synthetic users, and bridging the gap between model-level and product-level evaluation.