Lost In Delegation: How Effective is Agentic Delegation on Novel Information?
Abstract
LLMs have quickly evolved from models used in isolation to agentic systems in which an LLM operates external tools coordinates other LLMs to solve complex tasks. This agentic augmentation has led to impressive capability gains. However, what constitutes an optimal delegation setup is still not systematically understood. For instance, it is unclear how composite tasks should be decomposed, particularly when they involve post-cutoff knowledge unavailable to the orchestrator itself. In this work, we use purpose-built datasets to compare various delegation modes. We find that, perhaps surprisingly, delegation provides no significant accuracy gains over non-delegated setups. Furthermore, when tasks involve knowledge unavailable to the orchestrator, delegation can become highly inaccurate, while direct tool use with the same LLM performs much better. Our analysis identifies distorted delegated queries and the questioning or rejection of executor answers as potential failure modes at agent interfaces.