Beyond the verdict: Measuring effect of semantic factorization of safety judgment
Abstract
Current LLM-as-a-judge systems often output a single safety verdict, which obscures separate decisions regarding policy applicability, violations, severity, and behavior. We study the effect of explicit modeling and inferencing of these decisions on judge model behavior. Across both inference-time prompting and task-specific training, we compare holistic, joint multi-axis, and independent axis-specific judges. Independent axis-specific judges show overall better performance, for both prompting and training paradigms. Failure analysis shows the largest gain of decomposition from policy applicability errors. Robustness experiments further reveal that decomposition itself is not beneficial uniformly. Joint multi-axis training is highly robust meaning-preserving and policy-relevant perturbations, yet remain sensitive to the autoregressive ordering of its outputs.