Extractable Memorization from a Differentially Private Large Language Model
Nirav Diwan ⋅ Gang Wang ⋅ Daniel Alabi
Abstract
In September 2025, Google released VaultGemma, a 1B-parameter language model trained from scratch with differentially private stochastic gradient descent (DP-SGD). The accompanying technical report found 0\% detectable memorization under a discoverable-extraction test. We want to revisit this result, given the substantially higher extraction rates of comparable non-private models of comparable size. % We argue this number reflects the *audit*, not the *model*. VaultGemma's guarantee is *sequence-level* DP. It bounds the influence of any single training sequence, but it does not prevent a model from reproducing content duplicated \textit{across many} sequences --- precisely the content most likely to be extracted. Because the report audits by sampling training data uniformly, it almost never draws such sequences and therefore systematically under-reports worst-case risk. We test this directly. On 14,000 well-specified, frequent, non-trivial prefix--suffix pairs from the Pile extraction benchmark, VaultGemma exhibits 7.6\% exact and 12.7\% approximate discoverable memorization. This is substantially far from zero, and only $\sim$30\% below models of similar size or trained on a similar recipe. % We further run an untargeted probe with simple PII-oriented templates: even within a budget of 200 queries, 1\% of completions yield externally verifiable personal information. % Our findings show that empirical assessments of DP-trained LLMs are only as strong as the sequences they audit, and motivate evaluation definitions that fix a *worst-case* sampling target rather than a *uniform* one an evaluator which may provide an incomplete picture.
Chat is not available.
Successful Page Load