Mos'ka: A Simple and Small Agent That Orchestrates Small and Not-So-Small Agents
Abstract
How small can an effective multi-agent orchestrator be? We introduce \emph{Mos'ka}, a small agent that solves tasks exclusively by orchestrating subagents through a simple yet expressive two-tool harness. Delegation may compensate for its limited model capacity, allowing a small orchestrator to manage both small and larger subagents. The \texttt{spawn_subagent} tool starts a fresh subagent in the background, while \texttt{wait} lets the orchestrator sleep until a new report becomes available. Decoupling spawning from waiting allows Mos'ka to act on individual reports and launch newly unblocked subtasks while other subagents continue running. It then synthesizes the final response from subagent reports. Within this fixed harness, we elicit an orchestration policy from a large teacher model through a detailed system prompt. The policy delegates all substantive task work to subagents, limiting the orchestrator's role to managing their work and communicating their results. It compares decompositions by expected critical-path latency, greedily schedules unblocked subtasks, and adapts the decomposition as reports arrive. We distill this policy into a small student through supervised fine-tuning on successful teacher orchestration traces, keeping subagent models frozen. On held-out Workplace Assistant tasks, the student achieves pass@1 within 3.2 percentage points of its teacher across three worker sizes, using roughly half as many orchestrator tokens. With E2B workers, it raises pass@1 from 13.6\% for the matched direct agent to 55.1\%. With 26B-A4B workers, pass@1 remains unchanged, while pass@3 improves. Transfer to two GAIA2 splits, excluded from orchestration training, is uneven: with 26B-A4B workers, Mos’ka improves Search pass@1 from 22.9\% to 29.0\%, but reduces Execution pass@1 from 22.3\% to 11.0\%. These results demonstrate that a small distilled orchestrator can approach teacher performance in a training environment, while its benefits depend on worker capability, task type, and the success metric.