LDPCache: Locally Differentially Private Multi-Query Processing with Cache Optimization for Large Language Models
Abstract
The Model-as-a-Service (MaaS) paradigm enables resource-constrained edge users to access cloud-based large language model (LLM) services but raises significant privacy concerns. Existing approaches perturb user queries under local differential privacy (LDP) before sending them to LLMs. However, naively applying these methods to multi-query scenarios leads to linear growth in privacy budget consumption. To address this, we propose LDPCache, the first cache-enhanced LDP framework for multi-query LLM processing. By leveraging semantic correlations across consecutive user queries, LDPCache maintains a cache of historical responses and intelligently decides whether to reuse a cached result or query the cloud-based LLM, thereby reducing privacy budget consumption. The core of LDPCache is an LDP-compliant scheduler that evaluates cache reusability based on both semantic similarity and the expected accuracy gains from LLM access. This scheduler enables hit decisions without actual LLM access while ensuring that the scheduling process itself satisfies LDP guarantees. To accommodate the limited memory of edge devices, we further design a multi-factor cache eviction strategy that balances query diversity, response accuracy, and recency to improve the hit ratio. Extensive experiments on real-world datasets demonstrate that LDPCache significantly improves response accuracy over baseline methods.