Toward Adaptive Safety Policy Enforcement for Self-Harm Risk in Conversations with Minors
Arian Abbasi ⋅ Ted Kwartler ⋅ Alan Aqrawi
Abstract
Products that deploy conversational AI for or around minors face a maintenance gap: youth language, trends, and probing behavior can shift quickly, while retraining, blocklists, and hand-tuned guardrails update over weeks to months. We present a human-supervised framework that semi-automates the maintenance of guardrail policy text, applied here to a proxy for under-18 suicide and self-harm (U18-SSH) risk. To exercise its components we build a synthetic allergen-severity corpus of 812 multi-turn conversations, and over it an embedding map (conversation-level embeddings, UMAP, HDBSCAN) that a reviewer opens to read concentrated regions before writing policy. We then compare five automatic prompt-optimization systems, one our own (APO+memory), across three policy-conditioned judges and two proposer sizes. The corpus's three risk tiers (g2 informational, g3 personal risk, g4 emergency in progress) form a laddered analog to the Columbia-Suicide Severity Rating Scale's (C-SSRS) severity-of-ideation subscale, a design referent. Three briefed raters reproduced the tier rubric on a 48-item pilot at Fleiss' $\kappa$ = 0.979. We optimize policy text on one split and score it on a separate 321-item evaluation set: APO+memory reaches the highest observed mean macro-F1, 0.955, against the hand-written seed's 0.449. A label-free UMAP and HDBSCAN fit isolates the g4 emergency tier at 68 of 69 items (98.6% post-hoc purity), while informational g2 stays mixed into the normal cluster. We contribute the proxy corpus to the community. It contains no real U18-SSH disclosure and its tiers are generator-defined rather than clinically validated; results are exploratory and do not measure child-safety performance.
Chat is not available.
Successful Page Load