EngramState: Loadable Tool Priors for Efficient Function Calling
Abstract
Large language model (LLM) agents typically perform function calling by retrieving relevant tools and reinserting their specifications as textual prompts at inference time. Although retrieval reduces the number of candidate tools, selected tool specifications must still be repeatedly re-encoded, incurring substantial latency and memory overhead, especially in on-device environments. We argue that this inefficiency stems from treating tool knowledge as transient text rather than reusable execution memory. To address this, we propose EngramState, a state-centric framework that compiles tool specifications offline into reusable recurrent state priors and directly loads them at runtime for query-only inference. EngramState combines retention-aware state construction, channel-wise state compression, a lightweight same-backbone state retriever, and feedback-driven state routing. On the DroidCall benchmark with an RWKV-7 1.5B backbone, EngramState reduces prompt tokens from 985 to 22 (97.8\%) under the Top-4 retrieval setting and substantially lowers time-to-first-token (TTFT) on a Galaxy S25 Ultra while maintaining or improving function-calling accuracy. Furthermore, the proposed state retriever achieves competitive retrieval quality using only 8.24\,MiB of retrieval storage. These results demonstrate that reusable recurrent states can serve as an efficient execution-time memory interface for scalable on-device function calling.