The Child in the Loss Mask: Children’s Reactions in Fine-Tuning Data Are an Unreported Safety Decision
Abstract
Child safety in AI concentrates, rightly, on what models say to children. We argue the field must also attend to the opposite direction: children are entering training pipelines as users and consequently as evaluators of model responses. Children may respond to model generations with affects over contents and behave inconsistently instead of consistently, providing weak assessment and contingency towards model outputs and their quality. As a result, any pipeline that fine-tunes on children's transcripts could impact downstream model behaviors incidentally. Moreover, the pipeline might have different unreported loss-masking conventions, which decide whether children's turns are trained on or merely conditioned on. Drawing on our controlled study of in-episode feedback and emergent misalignment, we defend three positions: (1) the loss mask over user turns is a child-safety hyperparameter. With byte-identical data, the same reactions raise or lower misalignment depending on mask location and model family; (2) what training on reactions chiefly produces is a predictive model of the evaluator, the substrate of potential manipulation and engagement-driven retention of children — a capability that behavioral evaluations do not see; and (3) context-route misalignment mitigations only gate misalignment rather than removing it, so plain-prompt evaluations do not detect the gated capability that a trivial cue instead restores. We close with recommendations for reporting, auditing, and proxy-based evaluations, and with open questions for the workshop.