Route, Don’t Overload: Large Tool Spaces Are a Cost Problem, and the Payoff of Subagents Depends on the Model
Abdul Matin ⋅ Fatima Taheri Dezaki ⋅ Seyed S Saboksayr ⋅ Naveed Ahmed Janvekar
Abstract
LLM agents now reach tool spaces of hundreds of MCP servers and thousands of APIs, and the common design of placing every tool schema in one agent's context scales poorly. Two fixes are used in practice but rarely compared under controlled conditions: retrieving a small set of tools into a single agent, and splitting the tools across an orchestrator and several subagents. Multi-agent frameworks claim that giving each agent fewer tools improves success, yet we know of no study that treats the tool-allocation policy itself as the variable under test, and multi-agent results are seldom reported per unit of token cost. We build a controlled $\tau$-bench testbed that holds the base model fixed and grows the tool space with injected distractor tools, and we compare six allocation policies: all tools in one agent (\textsc{S-All}); retrieval into one agent (\textsc{S-RAG}); a single agent that re-retrieves its tools each turn (\textsc{S-Dyn}); and three multi-agent policies whose subagents receive, respectively, dynamically retrieved neighbor tools (\textsc{M-Neighbor}), static similarity-based clusters (\textsc{M-Domain}), or dependency-aware clusters built from tool co-invocation structure (\textsc{M-Dep}). Our headline metric is reliability per token ($\mathrm{pass}^k$ per 1{,}000 tokens), reported alongside raw success, as we sweep the tool space from the native size to 120 tools across the retail and airline domains, and out to 400 tools on retail. The whole pipeline uses open weights and no proprietary dependency: DeepSeek-V3.1 is the base model for both the agent and the simulated user, and all-MiniLM-L6-v2 is the retriever. To probe robustness we also test two further open base models of comparable scale, a weaker open base model, and a harder semantic-distractor regime, and we correlate the per-task single- vs.\ multi-agent gap with three task-structure features (required-tool-set size, trajectory length, and dependency depth). Our central claim is that for a capable model a large tool space is a cost problem, not an accuracy problem, and that the cost gap of a naive single agent comes from re-sending every tool schema each turn rather than from single-agent design as such. Giving one agent all tools inflates cost $\sim$$4\times$ with no significant accuracy change, and dynamic query-conditioned selection removes that cost; but a single dynamic agent, while keeping cost flat, loses accuracy as the space grows. Partitioning tools across subagents with per-subtask retrieval gives the best accuracy at large tool spaces at a fraction of the naive cost; a single dynamic agent is competitive on the easier domain, but only the multi-agent policies stay robust on the harder, longer-horizon one. The cost balloon is universal across three open base models (DeepSeek-V3.1, Kimi-K2.5, Qwen3-235B), but the best remedy is model-dependent: split tools per subtask for a strong coordinator, use a single dynamic agent for a weaker one.
Chat is not available.
Successful Page Load