Executable Contracts for Gene-Set Enrichment: A Controlled Fault-Injection Study
Abstract
Agentic systems now write and run gene-set enrichment analyses end to end. We ask a narrow, measurable question: how far do deterministic, executable pre- and post-conditions on an analysis specification go towards preventing incorrect specifications, and where do they fail? We build an over-representation-analysis harness over Reactome and Gene Ontology gene sets, inject eight families of specification faults via controlled fault injection, and evaluate a suite of five guards across 32,200 executions with fixed seeds. On a held-out evaluation block, a fixed-threshold guard suite plus deterministic repair cuts synthetic task error from 12.6% to 1.8% (paired bootstrap over 600 base-query clusters, a reduction of 10.8 percentage points, 95% CI 8.8--12.6 percentage points) at full coverage. Guard time is 147 ms versus 3002 ms for the analysis under contention. Three negative results are notable: spreadsheet date coercion causes no measurable error increase for 150-gene queries; background padding severity is highly feature-space dependent; and when a dataset manifest lacks an annotation release, the provenance guard has a complete blind spot, leaving outdated-annotation faults undetected. Guards are only as good as the metadata they read. We release code, pinned data checksums and all machine-readable results.