Beyond Success Rates: A Risk-Adjusted Framework for Evaluating LLM Agents
Cathy Li
Abstract
Agentic AI is entering the life sciences---literature copilots, bioinformatics pipelines, and the first closed-loop laboratories---yet these systems are still evaluated one scalar at a time: capability benchmarks report task success rates and safety benchmarks report attack success rates---almost never both for the same system. We formalize a joint, risk-adjusted view: a risk-adjusted score $RAS_{\lambda} = E[X] - \lambda E[R]$ that trades capability against risk exposure at an explicit, deployment-dependent tolerance $\lambda$; a capability--risk Pareto frontier whose dominance judgments are provably invariant to $\lambda$; and a measurement protocol that co-reports both axes on the same trajectories across an attacker-budget sweep. Because no life-science benchmark yet reports the two axes as a joint frontier with an explicit tolerance, we validate the framework by deterministic reanalysis of published, paired measurements spanning three model generations---the AgentDojo 2024 paper cohort, its post-publication results-database cohort (mid-2024 to early 2025), and the 2025--2026 frontier (Claude 4.5/4.6, GPT-5.2, Gemini 3 Pro). The joint lens changes what the record says: the capability--risk coupling is era-dependent ($r = +0.73$, then $-0.23$, then $-0.94$), so ``inverse scaling of safety'' is a property of a training era, not a law; 14 of 17 evaluated models are Pareto-dominated; 2024 rankings flip at $\lambda^\star \approx 0.51$ while the post-publication cohort is $\lambda$-robust to 1.50; one safety refresh cut attack success by 32.8 points at zero utility cost ($z = 15.3$); and adaptive attackers inflate single-shot attack success by up to 75$\times$ at $k=200$. We instantiate the framework for life-science deployment---a biosafety-committee reading of $\lambda$ and a worked tool-permission example whose tolerance threshold sits at $\lambda^\star = 0.59$---and distill seven reporting guidelines for the next generation of agentic benchmarks, biosecurity-aware evaluation included.
Chat is not available.
Successful Page Load