Teen-Informed Robustness Benchmarking for Hurtfulness Detection
Abstract
Content moderation systems trained on adult-generated and platform-scraped data often struggle with youth evasive language (e.g., “unalive”, “school bop”). We present a pilot study through a five-day Young Research Assistantship Programme (YRAP), where 12 adolescents (aged 14-16) co-research and develop obfuscation strategies. Seven participants generated 423 variations of 61 hurtful and not-hurtful sentences. We benchmarked 11 hurtful detection classifiers, including commercial APIs, character-level models, transformers, hybrid architectures, and a zero-shot large language model (LLM). Youth obfuscation reduced performance across all systems asymmetrically, with hurtful content more often misclassified as not-hurtful. Letter-to-digit, emoji substitutions, and youth phrases bypassed every system tested, including LLM safety refusals for self-harm-related text. These findings demonstrate the value of ethically grounded, youth-centred adversarial data curation to identify blind spots in content moderation systems. Data and code are available at https://github.com/emnlp-anonymous/teen-informed-benchmarking