TVD-Conflict: Disagreement-Guided Discovery of Jailbreak Judge Vulnerabilities
Abstract
Jailbreak research increasingly relies on LLM judges to determine whether a model response fulfills a malicious goal. Yet the same response can receive different verdicts across judges, making reported attack success sensitive to the evaluator. We ask whether such disagreements can be turned into an automatic signal for discovering judge vulnerabilities. We introduce TVD-Conflict, an agentic framework that searches for realistic query–response pairs on which two judges, operating under the same evaluation rubric, disagree. Their judgments provide iterative feedback for revising each candidate, allowing the framework to collect difficult cases without requiring human labels during the search. Experiments across ten jailbreak attacks, four victim models, three judge backbones, and two established evaluation protocols show that evaluator choice can substantially alter measured attack success. Across 50 malicious goals, TVD-Conflict discovers 1,702 conflict pairs that expose recurring vulnerabilities in how judges interpret rubrics, respond to stylistic and refusal cues, reason about obfuscated content, and ground their verdicts in evidence. These findings show that disagreement is not merely evaluation noise, but a scalable signal for discovering jailbreak judge vulnerabilities.