MOSAIC: Scaling Long-Horizon Language Agents via Multi-Scale Adaptive Inference Control
Abstract
Large language model (LLM) agents have shown strong potential in tackling complex, multi-step tasks, yet scaling them to long-horizon settings remains fundamentally constrained by two problems: context window saturation and uniform allocation of inference-time compute across steps of vastly different cognitive demands. We introduce MOSAIC (Multi-Scale Orchestrated Scale-Adaptive Inference Control), a hierarchical agent framework that decomposes long-horizon decision making into three temporal scales: strategic, tactical, and operational, each governed by distinct context scopes and reasoning budgets. Central to MOSAIC is a Scale Scheduler, a lightweight policy trained via reinforcement learning that dynamically determines (i) which abstraction level should be invoked at each time step, and (ii) how many reasoning tokens to allocate, based on the estimated decision complexity. We further propose inter-scale context distillation, a bidirectional information compression mechanism that maintains coherent state representations across scales without redundant token consumption. Theoretically, we show that under a fixed total inference budget, MOSAIC's adaptive allocation policy achieves a strictly lower regret bound than any fixed-rate policy in a hierarchical semi-MDP formulation. Empirically, MOSAIC achieves state-of-the-art or competitive results on four diverse long-horizon benchmarks: SWE-Bench Verified, BrowseComp, WebArena, and GAIA (Level 2&3), while reducing total inference tokens by 38--57\% relative to flat ReAct baselines. Ablation studies confirm that each component of multi-scale decomposition, adaptive scheduling, and context distillation contributes meaningfully to both performance and efficiency.