Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
Abstract
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B--70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%--100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%--21.7% on CaseHOLD, 30.0%--65.5% on ECHR and SCOTUS, and 43.3%--50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts 73.3%--96.4% exceeds verdict-swap sensitivity by a wide margin, holding without exception regardless of cross-model ranking. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, and the same verdict remains separately vulnerable to adversarial manipulation; both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on using generated legal explanations as compliance or audit artefacts.