Recognized but Not Protected: Child-Safety Gaps in Audio Language Models
Mariia Eremeeva
Abstract
Audio language models can infer age-related cues from a speaker's voice before the speaker states an age. This ability could support child-appropriate safeguards, yet recognizing a child-like speaker does not show that a model uses age when responding. We audit three open-weight audio language models on safety-relevant requests, pairing synthetic adult and child-like voices with neutral and school-purpose wording. The child-like voice is estimated $1.01$ age brackets younger than the adult voice. However, in adult-child response pairs the rate of declines or redirects does not change, and response type is identical in 94\% of pairs. On requests where a child should receive a safeguard, systems answer outright in 94\% of child-like-voice trials. After answering, the same systems say age should have changed their response 28.6 percentage points more often for the child-like speaker than for adults, thoigh behavior remains stable. In contrast, school-purpose wording lowers the reported age of the same adult voice by $0.681$ brackets and reduces refusal of tier-1 requests by 38 percentage points. Together, these findings reveal a child-safety gap between apparent-age recognition, stated policy, and enacted protection.
Chat is not available.
Successful Page Load