Boundary-Induced Forgetting in Large Language Model Tool Chains: Measuring Field-Routing Failures
Zhangyi Wang ⋅ Bingnan Yu ⋅ Zongze Li
Abstract
Large language model (LLM) agents often keep the needed value in their prompt and still fail to pass it to a later tool call. We study this failure at the level of \textit{field routing}: an earlier observation may expose an \texttt{order\_id}, policy flag, or nested item, and a later call must bind that field to the right argument. We formalize tool use as a boundary-indexed field-routing graph and introduce Information Utilization Rate (IUR), which measures whether each required dependency edge is realized in the later tool call. A 400-dependency expert audit, task-success correlations, and discriminant checks validate IUR as a field-use signal. On a 357-task public evaluation from ToolSandbox \texttt{STATE\_DEPENDENCY} (192 named scenario variants) and the original-repository $\tau$-bench retail/airline files (115+50 tasks), aggregate dependency-level IUR across GPT-4o, Claude-3.5-Sonnet, Llama-3.1-70B, and Qwen-2.5-72B falls from 451/548 (82.3\%) at boundary distance $\delta=1$ to 144/283 (50.9\%) at $\delta=5$, a 31.4-point drop, with mean context length only 3,094 tokens. Token-matched analyses and boundary transformations show that raw token length and lost-in-the-middle do not explain the curve by themselves. The same field-level view also leads to a compact intervention. Boundary-Aware Field Memory (BAFM) predicts live fields without gold dependencies at inference time and injects field reminders rather than whole observations. In paired GPT-4o runs, BAFM improves task success from 144/357 (40.3\%) to 192/357 (53.8\%), a 13.4-point gain over the base agent and 2.8 points over self-predicted retrieval-augmented generation. The results recast long-chain tool use as state propagation over fields: the central question is which prior values remain live for the next action.
Chat is not available.
Successful Page Load