When Accuracy Is Not Enough: Evaluating Trustworthiness in High-Stakes Lease Interpretation
Abstract
Large language models are increasingly used to explain legal and contractual information, but standard evaluations centred on answer accuracy can miss failures that are consequential in high-stakes settings. In lease interpretation, a response may appear correct while relying on inapplicable legal material, unsupported evidence, or an inappropriate decision to answer. We present an evaluation framework that characterises trustworthiness along three dimensions: legal correctness, evidence integrity, and decision safety, together with explicit hard-failure criteria for serious errors. We study this framework in New South Wales residential tenancy using a 2,000-question benchmark and a stepwise comparison of progressively safeguarded system configurations, from direct generation to applicability-aware retrieval, evidence-first generation, evidence checking, and selective response behaviour. On the protected test set, answer accuracy increases only from 82.3\% to 84.0\%, while the Trustworthy Response Rate (TRR) rises from 58.7\% to 82.3\% and the Hard Failure Rate (HFR) falls from 20.7\% to 5.7\%. Two configurations with identical 83.7\% Answer Accuracy achieve different TRR (72.0\% versus 78.0\%) and HFR (11.7\% versus 8.0\%). These results show that configurations that appear equivalent under answer accuracy can differ substantially under response level trustworthiness measures.