Can an LLM Reason Like a Lawyer? Benchmarking the ability of LLMs to map the facts of a case to the elements of the applicable legal rule
Abstract
Legal judgment prediction has made substantial progress in recent years, yet existing work evaluates whether models reach correct conclusions rather than whether they reason correctly. The core of legal decision-making is the mapping of case facts to the elements of an applicable legal provision, a step that prior benchmarks neglect. We introduce the first benchmark for evaluating legal reasoning, built on 375 judgments of the European Court of Human Rights (ECtHR) discussing Art. 11 ECHR on freedom of assembly. We exploit that the court routinely splits the decision in a stage reporting the facts of the case, and an evaluation stage in which it uses these facts selectively for justification. We give the (preprocessed) list of facts to 10 downstream LLMs and evaluate their justifications. The pipeline is validated by three expert legal annotators. Downstream models are reasonably accurate in predicting whether Art. 11 has been violated, and which elements of doctrine are critical, but perform substantially worse at attributing specific facts to doctrinal elements. Essentially they make the right decision, but for the wrong reasons. This weakness is invisible under standard evaluation metrics. Our benchmark, pipeline, and evaluation framework are released. They can easily be adapted to any legal rule that follows the Civil law tradition, i.e. to legal reasoning that requires the mapping of facts to an abstract legal rule.