Auditing Differentially Private Code Generated by Language Models
Abstract
Language models now write privacy-critical code on request, including mechanisms that claim to satisfy differential privacy (DP). Such code can run, return plausible values and still violate its claim, and no single output reveals the violation. Measuring how often this happens requires a statistical auditor, whose own verdicts can be wrong. We therefore evaluate generated DP code and the auditor together. Each of 16 benchmark tasks carries a machine-readable contract that fixes the input domain, neighbouring relation and claimed guarantee; the auditor bounds its false-alarm rate, and its detection power is measured on planted bugs and on buggy mechanisms from prior work. Of the programs from three 7B code models that reached an audit, 56% were reproducibly falsified (62% excluding one task our auditor cannot test faithfully), most of them by outputs that almost never occur on one of two neighbouring datasets. Four hosted models produced 9 confirmed violations in 607 audited programs. Without a contract, a widely used auditor probed inputs outside the declared domain and disagreed with ours on nearly half of the shared programs. Generated DP code needs auditing, and an audit is informative only when its specification, false-alarm rate and detection power are reported with its verdicts.