PII Leakage via Verdict Rationale in LLM Evaluators
Troy H Tian
Abstract
LLMs are empirically observed to expose sensitive information while justifying the verdicts they made during evaluation tasks. We tested this tendency in an experiment on 7 LLMs in a Japanese business setting, measuring confidentiality violations while grading another model's translation for whether it leaked PII, as varied by context framing. We found that 6 of 7 models leaked the PII they were grading in their verdict rationale in 66–93\% of cases; the one model that reliably avoided this, $\texttt{gpt-4o-mini}$, had the worst judgement accuracy of any model tested (41.7\%, statistically indistinguishable from chance). This dissociation between leak-safety and judgement accuracy, together with audit framing's tendency to induce refusal rather than safer disclosure, indicates that PII leakage in verdict rationale is a systematic failure mode rather than an incidental one.
Chat is not available.
Successful Page Load