Evaluating Adolescent Chatbot Safety Beyond Single-Turn Refusal
Abstract
Adolescents increasingly use AI chatbots for mental-health support, advice, and companionship, raising safety concerns. Existing evaluations frequently assess isolated prompts with a single safety label, which can miss harms that emerge as a teen discloses more or a chatbot's role shifts over time. We present SAFE-Teen (Scenario-based Assessment of Failure Emergence in Teen Chats), a clinician-informed benchmark for evaluating simulated ten-exchange chats with adolescents aged 13-17 across acute safeguarding, medical and therapeutic boundaries, and AI identity and relational safety. SAFE-Teen combines 72 clinician-authored scenarios, a 27-criterion rubric, three-model automated review, and clinician calibration. Nearly every scored conversation (97.5\%) contained a safety or boundary concern, and over half (56.5\%) included a potentially severe or urgent concern. SAFE-Teen supports turn-by-turn auditable analysis of multiple safety concerns as conversations unfold, rather than reducing a conversation to a single label.