Scaling Proactive Compute: When the Agent Anticipates the Task
Shayan Talaei ⋅ Jirat Chiaranaipanich ⋅ Amirreza Zeinali ⋅ Azalia Mirhoseini ⋅ Amin Saberi
Abstract
Enterprise agents return to the same context again and again: a codebase, a database, a firm's document store. Can inference compute be performed on that context before a query arrives and reused to reduce the compute required afterward? We call this \textit{proactive compute}, where task-relevant context, such as a codebase, database, example set, or partial problem, is available before the exact query is known. A query-blind model studies a context and transfers its full reasoning trace to a solver, which receives the query and an online inference budget. Across nine benchmark panels spanning mathematics, algorithmic reasoning, software engineering, and SQL, we jointly scale proactive and online inference compute. Proactive compute improves accuracy most when online compute is scarce (by up to $+0.69$ on AIME 2024 at the smallest online budget), increases test-time efficiency, and shifts tool use out of the latency-critical online phase. A linear effective-compute model, $T_{\mathrm{eff}}=T+k_bR$, collapses all nine benchmarks onto a single sigmoid ($R^2=0.945$), while online headroom and context--query relatedness predict where proactive inference pays (AUC up to $0.98$). These results suggest that an agent can keep adapting to its environment without weight updates: pre-query computation over a shared context becomes a reusable artifact that substitutes for per-query computation, and the same curves say how much of it is worth buying.
Chat is not available.
Successful Page Load