DUO: Practical LLM Unlearning for Child Safety via Data and Optimization Co-Design
Abstract
LLM unlearning aims to remove hazardous or privacy-sensitive information without retraining from scratch. Practical unlearning, however, must simultaneously remove undesirable behavior, preserve general utility, avoid over-refusal on neighboring concepts, and remain robust to adversarial probing. These demands are particularly acute for child-facing LLMs, where unsafe requests can closely resemble legitimate safety questions. We introduce KidBoundary, a benchmark of 230 semantic triplets (690 prompts) across five risk categories, where each triplet pairs an unsafe target with a benign neighbor and an adversarial reformulation. We also propose DUO, which unifies heterogeneous objectives as QA data, sharpens their boundaries with contrastive anchors, and applies intention-guided bidirectional Top-K logit distillation. DUO reduces unsafe compliance and attack success, achieves the best triplet-level Boundary Success, and preserves benign response quality and general utility. Results on MUSE-Book and WMDP-Cyber further demonstrate generalization beyond child safety.