From Evidence to Intervention: Verification-First Agents for Scientific Hypothesis Generation
Abstract
AI-scientist systems can draft plausible scientific answers faster than experts can review them. In experimental science, physical confirmation is often slow and available records are incomplete. The immediate question is whether an answer's evidence is strong enough for delivery before it guides a laboratory decision. We study this question in microwave plasma chemical vapor deposition (MPCVD) single-crystal diamond growth. Our workflow routes each question to evidence, checks database queries, drafts a cited answer, compares each claim with its cited evidence, and revises failures. Here, verification is a delivery decision based on traceable evidence, not a certificate of scientific truth. A claim passes only when the evidence supports every material statement. Unsupported or contradicted claims receive targeted revision; persistent failures are removed or qualified. Without verification, 77% of test answers would reach the user with at least one unsupported or contradicted claim. With verification and revision, this rate falls to 10%, an 87% relative reduction. The system correctly identifies whether a test answer requires revision in 84.0% [75.6%, 89.9%] of cases and detects 95.5% of individual claims that require revision. For literature questions, the top five results contain a relevant passage for 70.0% [56.2%, 80.9%] of queries, and 69.9% of all retrieved top-five passages are relevant. For structured-data questions, we observe no numeric fabrication by the database agent. Although conservative, the verifier combined with targeted revision restores full-answer delivery to approximately 90% without increasing the observed violation rate. These results show that an imperfect verifier can be useful when evidence checking is coupled to repair rather than final rejection.