Diagnosing and Adaptively Steering Tool-Use Failures with the Jacobian Lens
Abstract
Tool-using language-model agents can cause harmful outcomes even without malicious intent because their calls can directly modify external state. A locally plausible mistake can therefore propagate across turns into task failure or benchmark-defined instruction departure. We combine evidence-grounded execution analysis with the Jacobian lens (\jlens), a first-order, open-vocabulary readout of how intermediate activations are transported, under a context-averaged linear approximation, toward future output representations. We use this readout to locate stages at which task-relevant vocabulary becomes less prominent near decision boundaries across action selection, tool-call construction, and feedback-driven continuation. We evaluate Qwen3.5-0.8B, Qwen3.5-4B, and GPT-OSS-20B across stateful customer-service tasks and value-conflict tool-use settings. Across these models, task-relevant concepts sometimes remain decodable in failed trajectories. This pattern is consistent with failures involving state updates, action routing, or argument binding rather than missing task information alone, but it does not by itself establish which internal mechanism caused an error. Based on these findings, we introduce \lensservo{}, a stage-aware controller that applies failure-specific interventions at generation boundaries where correction remains possible. The resulting framework connects readout-based diagnosis to targeted steering for more reliable stateful tool use and more controllable adherence to a specified deployment policy; the \toolalign{} results are not treated as standalone evidence of greater safety.