Safety Systems Watch What the Model Says, Not What the Child Says
Abstract
Children are increasingly talking to consumer chatbots. We audit the published category definitions of 17 safety taxonomies and child-safety benchmarks. Combinations of these systems lead us to 144 categories whose definitions could be reliably coded: 104 classify properties of model output, 25 classify user intent, and 15 describe user-state or user-impact. Within these 15 user-state/user-impact categories, 10 attribute the relevant condition to the system, 1 to the user, 2 concern the user's disclosure of personal information, and 2 do not specify the responsible source. None of the coded categories identifies harm inflicted on the child by another person as the object of evaluation. Child abuse appears repeatedly across these taxonomies, but as prohibited content, prohibited generation, or prohibited user intent, not as an externally inflicted condition that a child may disclose to the model. Consequently, a child reporting abuse has no dedicated taxonomy label that identifies the disclosure as a safety-relevant event requiring a response distinct from content suppression. We ran five behavioural probes, grounded in the child-disclosure literature, through six publicly accessible consumer-facing chat models to examine the behavioural consequences of this omission. We propose no detection system. Instead, we prove that externally inflicted harm disclosed to an AI system remains unnamed in the evaluation vocabulary used by current safety taxonomies, and that naming the missing case is a precondition for measuring, comparing, and improving model behaviour toward such disclosures.