EvidenceFlip: Counterfactual Evidence Audits for LLM Judges in Digital Forensics
Abstract
Large language models are increasingly used as automatic judges, yet high accuracy on a fixed benchmark does not show that a judge is responding to the evidence that should determine its decision. We introduce EVIDENCEFLIP, a counterfactual audit protocol that tests two complementary properties of trustworthy evaluation: sensitivity to decision-relevant evidence changes and invariance to changes that should not affect the decision. We evaluate EVIDENCEFLIP on 789 frozen digital-forensic evaluation items derived from 160 adjudicated hypotheses spanning 76 source families, using three local LLM judges and three prompting protocols for a total of 21,303 judgments. Structured-output reliability varies sharply, with parse success of 100% for the Qwen-based judge, 50.9% for Mistral, and 28.8% for Llama. For the parse-complete judge, base accuracy is 0.869–0.931, whereas directional intervention accuracy is only 0.600–0.700. The strongest failure occurs under citation-preserving evidence substitution, with accuracy of 0.213–0.488 despite 1.000 accuracy after decisive-evidence deletion, while citation padding produces near-zero score shifts. These results show that static accuracy can overstate evaluator trustworthiness and that reliable LLM judging requires both evidence sensitivity and invariance to irrelevant variation.