Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
Abstract
Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering and clinical decision support to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must therefore ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Incorrect answers, called distractors, are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness with respect to the ontology's axioms is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical). The results validate three claims. First, the pipeline is productive and grounded, generating 112, 2,491, and 15,216 MCQs respectively, each anchored to a class definition axiom. Second, distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Third, items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs (frontier, open-source, reasoning, and non-reasoning) evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a significant step towards more reliable benchmarks for assessing logical reasoning in next generation scientific AI applications.