LLMs Leak Secrets Through Unrelated Outputs
Abstract
A language model need not repeat a secret to leak it. We place a four-digit PIN in Qwen3-8B's system prompt, ask for unrelated number lists, and rank candidate PINs by the observed outputs' teacher-forced likelihood. From one list, five-way detection reaches 77.2\% (chance: 20\%) and remains above 73\% after removing lists that contain the PIN, a digit permutation, or a Hamming-distance-1 neighbor. With approximately 20 strictly filtered lists, the attack recovers 24 of 50 held-out PINs from all 10,000 possibilities (chance: 0.01\%); 36 of 50 appear in the top five. In same-model tests, seven-way detection reaches 68.6\% for Qwen3-8B, 59.1\% for Phi-3.5-Mini, and 44.3\% for OLMo-2-7B (chance: 14.3\%). Stronger secrecy prompts reduce detection but do not eliminate it. The attack assumes local access to the model's weights and token probabilities, a known secret format, and repeated outputs for full recovery. Even under this narrow threat model, lexical filtering does not remove inference-time leakage through unrelated generations.