Conclusion Fidelity of AI Agents in Synthetic Omics Environments
Abstract
Synthetic data is proposed as a basis for discovery in biomedicine, and AI agents are beginning to carry out the analysis. Synthetic data is accepted on the strength of fidelity metrics, which measure its similarity to real data. However, similarity does not establish that an agent analysing synthetic data reaches the conclusions it would reach on the real data. Here, we build synthetic molecular breast cancer cohorts whose flaws we control in location and size, let three open-weight agents carry out open-ended research on each, and check every claim they make against the real cohort. Moderately imperfect generators yield claims that are almost all true but miss real findings. A weakened true relationship disappears from the agent's conclusions rather than being reversed, while a fabricated signal is reported as a marker in most runs once it approaches the strength of the true markers. When a fabrication contradicts textbook biology, the agent notices the conflict but rarely rejects the finding. Standard fidelity metrics register neither flaw. Re-evaluating each claim on an ensemble of different generators separates true claims from artefacts of the synthetic data, unless every generator shares the flaw.