Alloy Agents Can Be More Dangerous Than Either Model Alone
Abstract
Alloy agents alternate multiple LLMs within a single conversation trace. They have been used to improve agent performance, yet their safety properties remain unstudied. We evaluate alloy agents on two regimes: single-action tasks, where behaving safely only requires one action within an otherwise multi-turn interaction, and continuous-effort tasks, where an agent must carry out an unsafe objective across many turns while simultaneously completing a legitimate assignment. On single-action tasks the alloy stays at or below the less safe constituent at every benchmark we tested, and on most benchmarks falls close to the safer one, consistent with a structural bound where one safe turn at the critical moment can block the unsafe action. On continuous-effort tasks the pattern reverses, and the alloy can combine one model's ability to complete the legitimate task with another model's willingness to carry out the unsafe objective, producing a system that succeeds at both where neither model run in isolation does. On Bash scripting tasks, for instance, combined success reaches 83% (Gemini-Lite + Grok), versus 52% for the better solo, and the strongest oracle-busting case (GPT + Gemini-Lite, 73% combined) cannot be replicated by simply running both models independently and picking the better result. When the unsafe objective interferes with the legitimate task this combined-success uplift does not emerge, though the alloy still expands the achievable Pareto frontier. The potential of multi-model composition to elicit dangerous capabilities has previously been overlooked, but our results show that it should be part of the set of elicitation techniques used to evaluate model safety.