When Guardrails Redirect: Excluding a Requested Action Can Trigger Another
Abstract
Large language model (LLM) agents can change external systems through tool calls. To reduce infeasible calls, runtimes may use database state before generation to constrain a tool's argument schema. We find that when this constraint excludes the user-requested target, it can redirect the model to another executable but unrequested target, turning a rejection with no state change into a wrong-target modification. We measure this failure using matched runs that differ only in whether the requested target remains generatable, and we execute the selected calls with official tools. Across three agent benchmarks, we observe the same conditional transition. When the run in which the requested call remains generatable ends in an official rejection, its matched target-excluded run modifies another target. Across five model configurations, excluding the requested target from the constrained output space never reduces the wrong-target modification rate. Diagnostic experiments further show that removing decoder enforcement so that the requested call remains generatable and can be officially rejected reduces the paired risk difference from 93.8 to 4.6 percentage points. The same final call can follow either a direct request or a substitution after the original target is excluded. A current-state check cannot distinguish these histories or determine whether prior authorization covers the action. We therefore evaluate request-aware execution checks that retain the original request, allowed alternatives, and approval for the final action. Strict request matching prevents all wrong-target modifications in the execution replay. In a separate proposal-level test, a set-based variant retains specified alternatives and rejects unlisted proposals. Thus, filtering generation by current state can change the target on which the model acts. Current-state validity does not imply consistency with the user's request.