Differentially Private Term Statistics for Efficient Synthetic Text Generation
Abstract
Large language models (LLMs) can generate differentially private synthetic text without private training. Current inference-only methods, however, often repeatedly generate, score, and refine candidates, requiring many LLM inference or API calls. We explore a more compute efficient alternative and ask: what minimal private representation of a sensitive corpus is needed to prompt an LLM to generate useful synthetic text? We propose to release first-order term-presence statistics under differential privacy and then sample prompt keywords from the release; for sequence-sensitive tasks, we also study an optional second-order extension. We show that this approach is competitive with state-of-the-art methods on sentiment and topic classification as well as next-word prediction tasks, while requiring far fewer LLM inference calls. We argue that a one-time DP release stage based on simple text statistics is therefore a strong baseline before using iterative private search. Code available at: \url{https://doi.org/10.5281/zenodo.22303113}.