From Facts to Personas: Interpretable Role Unlearning in LLMs via Mixture-of-Experts
Abstract
LLMs are widely adopted in multi-role dialogue and persona simulation. However, their strong ability to imitate role-specific linguistic and behavioral patterns can introduce privacy risks, reinforce social biases, and lead to unsafe or uncontrollable outputs. To address these issues, we introduce the Role-playing Unlearning task, which aims to selectively forget the style and knowledge associated with target roles while preserving the language ability and behaviors of the remaining roles. Unlike prior unlearning tasks that focus solely on removing factual information, our task additionally requires forgetting persona-specific characteristics, making the control over model generation more fine-grained. We propose ROLEMOE, a Mixture-of-Experts framework that enhances the separation of role-specific patterns by mapping different roles to more independent expert subspaces, providing structural interpretability. To reinforce this separation, ROLEMOE incorporates two complementary losses: a specialization loss that drives different roles to adopt different expert distributions, and a disentanglement loss that encourages different experts to develop distinct capability subspaces. This decomposed structure enables role unlearning to focus selectively on target-role experts. We finally propose RoleBench-Unlearn for evaluation, and experiments across multiple LLM architectures show that ROLEMOE achieves significantly more effective and controllable role forgetting, while better preserving the linguistic quality, knowledge, and persona consistency of non-target roles. All code and datasets are publicly available at: https://anonymous.4open.science/r/RoleMoE-6328.