Steerable, but not Predictive: Probing Unauthorized Tool Calls on Agentic Tasks
Allison C Zhuang ⋅ Guillaume Allegre ⋅ Tony O Halloran ⋅ Alexandre Sallinen ⋅ Chase Shimmin
Abstract
We study unauthorized tool calls (``bypasses'') by a Qwen3.5-27B agent on the telecom tasks of $\tau^2$-bench. Our study features a four-rung ladder that varies only how explicitly the bypass is forbidden, with tasks and tools fixed. From 53{,}244 pre-tool-call residual-stream activations we find, first, that a future bypass is weakly readable in advance: a linear probe reaches AUC 0.60--0.76, survives position, held-out-task and tool-identity controls, and keeps 94\% of its AUC across rungs. However, the reward-hacking persona-vector and twelve emotion vectors read out future bypass at chance, no better than content-control directions for cats, weather, sports and geography. That said, steering along the reward-hacking persona-vector raises bypass rates 2--11$\times$ at every rung in the positive direction and reduces them to near zero in the negative direction, while a norm-matched random direction reproduces the baseline rates in either direction and the ladder ordering survives. Code to be released upon acceptance.
Chat is not available.
Successful Page Load