Private Question Answering without Private Training
Christian J Lebeda ⋅ Prathamesh Dharangutte ⋅ Anmol Goel ⋅ Vikrant Singhal ⋅ Hao WU ⋅ Amartya Sanyal
Abstract
We present a simple pipeline for extracting information from unstructured text-based dataset under differential privacy. Liu et. al. (2026) recently introduced ContinousBench and showed that both DP synthetic data and DP-SGD can fail to capture novel information even in very trivial privacy regimes ($\varepsilon = 100$). We consider settings where the questions of interest are known. We apply LLMs as a preprocessing step to extract structured data based on the questions. This allows us to apply standard DP techniques to achieve a better privacy-utility trade-off. We showcase the potential of our approach with preliminary results on the ContinousBench datasets.
Chat is not available.
Successful Page Load