ReasoningShield: Safety Moderation over Reasoning Traces of Large Reasoning Models
Abstract
Large Reasoning Models (LRMs) leverage explicit reasoning traces, known as Chain-of-Thought (CoT) reasoning, to decompose complex problems into intermediate steps before deriving final answers. However, these reasoning traces introduce unique safety challenges: harmful content can be embedded in intermediate steps even when final answers appear benign. Our study reveals that existing state-of-the-art moderation tools experience significant performance degradation on CoT moderation, with F1 scores dropping by up to 35.2% compared to traditional answer moderation. To address these challenges, we present ReasoningShield, a comprehensive benchmark and a suite of strong lightweight models for CoT safety moderation. Our benchmark includes 9.2K annotated Query-CoT pairs with structured stepwise risk analysis, consisting of a 7K-sample training set and a 2.2K human-annotated test set, spanning 10 risk categories, 3 safety levels, and 8 LRMs with diverse reasoning paradigms. To further demonstrate the utility of the benchmark, we develop lightweight models with 1B and 3B parameters using a two-stage training strategy. Our models achieve 91.8% F1 on our benchmark, substantially outperforming leading tools such as LlamaGuard-4 by 35.6% and commercial models such as GPT-4o by 15.8%, while generalizing effectively across diverse reasoning paradigms and unseen scenarios. All resources, including the dataset, code, and models, are released at https://anonymous.4open.science/r/ReasoningShield.