Rethinking Hallucination in Large Language Models: A Legal Framework for Definition, Diagnosis, and Accountability
Abstract
Large language models (LLMs) are increasingly deployed in high-stakes domains such as law, where their ability to generate fluent but potentially incorrect outputs poses significant risks. A central challenge is the phenomenon commonly referred to as hallucination, which is typically defined in an underspecified manner that conflates factual errors, fabricated content, and reasoning failures, limiting its usefulness for diagnosis and mitigation. This paper argues that meaningful progress requires moving beyond a binary notion of correctness and instead characterizing why an output is incorrect. Drawing on legal reasoning as an analytical template, we propose a two-stage diagnostic framework that decomposes LLM errors into evidentiary grounding failures and inferential reasoning failures. We further clarify boundary cases that are often misclassified as hallucinations, including novel-but-correct outputs and controllability failures. We empirically validate the framework through three experiments on legal tasks using ChatGPT (GPT-5.2) and Gemini (2.5 Flash / 2.5 Pro), with expert legal practitioners involved in the evaluation process. Across diagnostic classification, targeted prompt-based mitigation, and external source grounding interventions, we show that the proposed decomposition enables reliable separation of failure types and that failure-type–matched interventions consistently outperform generic baselines. We further release a curated LegalAI Hallucination Benchmark of 100 queries across 9 legal task categories for evaluating LLM reliability in legal applications. Related resources are available at \url{https://github.com/Jhhuangkay/legal-llm-hallucination}.