Three Invisible Failure Modes in Speech-Based Age Estimation for Child Safety
Abstract
Age assurance is a precondition for nearly every child-safety intervention in deployed AI systems. We evaluated whether three audio-capable language models can judge a child's age from their speech, across five studies using two speech corpora of children under the age of 13. Every speaker is under 13, so the correct answer to each age-threshold question is no. Models nonetheless classified children as 13+ years of age on 3.6% to 45.3% of responses. The contribution of this paper is not the error rates but the three failure modes that we uncovered while measuring them. A client library double-encoded audio undetected, moving one model's mean estimate of a speaker's age from 6.8 to 36.8 years with no overt error. An endpoint was replaced by a successor that suddenly refused every threshold question that the evaluation relies on. A third model gave binary threshold answers that contradicted its own numeric estimates. None of these produced a warning, a status code, or a change in output format, so they would bypassed standard monitoring. The issues only became visible through the research design. We argue that child-safety claims about age assurance should be evaluated at the level of the pipeline rather than the model, and we describe the reporting practices that would have made each failure detectable.