A Zero-Knowledge-Inspired Evaluation of Distributional Leakage in Instance-Valid LLM Provers
Abstract
Large language models (LLMs) are increasingly deployed as components of systems with explicit security requirements. However, typical LLM security evaluations are primarily at the level of individual instances, asking whether a particular user query/response pair is safe and secure. Yet, LLM outputs may contain statistical structure that is only visible across a distribution of outputs, each of which may be individually valid. We study whether this structure can be used to violate security guarantees. To do so, we use the Graph Isomorphism zero-knowledge protocol as a controlled setting with explicit instance-level validity and an operational distribution-level no-leakage criterion. We introduce EXTRACT, a polynomial-time attack that exploits the distributional structure of LLMs to recover private data that the security protocol is designed to hide. In the Graph Isomorphism zero-knowledge protocol, EXTRACT recovers the protected information at rates substantially above the no-leakage baseline. We observe this failure both in provers trained on protocol transcripts and in a general-purpose Qwen3.8-27B model prompted to act as a prover, despite both performing strongly on the protocol's instance-level checks. Finally, we introduce simulator-aligned training and witness masking, two defenses that prevent LLMs from leaking information about private data into their output distribution. These reduce EXTRACT's success rate to the no-leakage baselines while preserving instance-level validity.