Traceable Is Not Supported: A Layered Metabolite–Disease Evidence Agent and What Its Validator Missed
Abstract
Scientific agents use large language models (LLMs) to search databases and summarise evidence. A common way to prevent hallucination is to check whether the retrieved records contain the names and numbers in the summary. But a value being traceable does not mean it is used correctly. We built MetaboCausal, a metabolite–disease evidence agent that searches for genetic evidence, effector genes, drugs and clinical trials. In the process, code builds the card, and an LLM writes only a short summary with the validator's check. This design blocked the most obvious hallucinations. It did not guess at inputs it could not read, and it blocked invented numbers. But in one test, a false nomination slipped past the validator: in a summary, GCSH and GLDC were nominated as "effector genes", although both genes came only from a metabolic-reaction database to add biochemical context. When we removed this database, the same summary was blocked. So adding an evidence layer made the validator weaker. Then we audited the validator with LLM agents' adversarial attacks in two rounds. In the first round, 51 test cases found eight kinds of defect, and seven of them held when tested again. The second round confirmed 16 more defects in three parts around the validator and rejected 12 candidate claims. We fixed these 16 defects, including the most serious one. Across the two rounds, four defects share one cause: the validator does not check how a sentence uses a value. Therefore, we recommend that validators check each value and how it is used in the sentence. Validators should also be audited again whenever an evidence layer is added.