LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
Abstract
Computer-use agents (CUAs) that interact with real systems can automate complex tasks, but introduce critical safety risks in long-horizon tool-use workflows. Many existing benchmarks rely on outcome-based evaluation, however, outcome-based evaluation cannot reflect agent's safety awareness, as an agent may reach a safe outcome by chance while taking unsafe steps during planning. To address this gap, we present \textsc{LPS-Bench}, which extends the outcome-centric evaluation towards LLM agents' safety awareness in tool usage (MCP/skill) over long-horizon tasks, as reflected at the action trajectory level. \textsc{LPS-Bench} comprises 570 test cases derived from 65 scenarios across 7 task domains and 9 planning-risk types, with user contexts spanning both benign ambiguity and adversarial steering. Results on \textsc{LPS-Bench} across 13 representative LLM agents reveal previously underexamined safety failures in long-horizon tool-use tasks, showing that current agents often struggle to maintain safety awareness under ambiguous benign instructions and adversarial steering. We further analyze key failure modes and evaluate lightweight mitigation strategies, showing that prompt-based interventions provide limited improvements yet remain insufficient for robust trajectory-level planning safety.