Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
Yen-Shan Chen ⋅ Zhi Rui Tam ⋅ Cheng-Kuang Wu ⋅ Yun-Nung (Vivian) Chen
Abstract
Current evaluations of LLM safety predominantly rely on *severity-based taxonomies* to assess the harmfulness of models' responses to malicious queries. We argue that this formulation requires re-examination as it implicitly equates model output with realized harm while neglecting *Execution Likelihood*—the conditional probability of a threat being realized by a real-world actor given a model's response. In this work, we introduce **Expected Harm**, a metric that weights the severity of a jailbreak by its execution likelihood, modeled as a function of *execution cost*. Through empirical analysis of state-of-the-art models, we reveal a systematic *Inverse Risk Calibration*: models disproportionately exhibit stronger refusal behaviors for low-likelihood (high-cost) threats while remaining vulnerable to high-likelihood (low-cost) queries. Mechanistically, we first provide an analysis using linear probing, revealing that while refusal mechanisms linearly separate queries by severity, they do not use execution likelihood as an indicator for refusal. Then, we use Olmo-3's open-source staged checkpoints to trace the root cause of miscalibration to the Supervised Fine Tuning (SFT) training data distribution (82.8\% of safety examples at cost levels 0–1) and show that cost representations are weak from the SFT stage onward. Finally, we demonstrate that this miscalibration creates a vulnerability: by exploiting this property, we increase the attack success rate of existing jailbreaks by up to $2\times$.
Chat is not available.
Successful Page Load