Assistive Variant Curation in Clinical Cancer Whole-Genome Sequencing: Evidence Ablation and Time-Sealed Evaluation
Abstract
We describe an assistive agent that curates driver variants in clinical cancer whole-genome sequencing at a two-site diagnostic laboratory and examine how evidence affects its adjudication. Five stages run ahead of a consensus conference of four board-certified reviewers, which finalizes every case and provides the reference labels used here. Two of the five stages call one frontier general-purpose multimodal model, in three calls, and no biology-specialized model is used at any stage; two more called a model earlier in development and now run on rules. On a development benchmark of 1,402 candidates in 40 cells, explicit predicates close 941 before any model is called, leaving 461 for model adjudication. Withholding read-alignment and copy-number images raises the share returned to a reviewer without a verdict from 28.6% to 42.3%, while withholding retrieval over past interpretations lowers specificity from 69.9% to 53.6%. Withholding the purity estimate supplied by the copy-number stage moves no measure by more than 3.0 percentage points. Yet accuracy pooled over the whole 1,402-candidate cohort remains within 94.2–94.8% across all arms, indicating that evidence can change which candidates the agent closes without a comparable change in the figure a concordance study would report. These measurements cover the verdict axis of reportability rather than the clinical-significance assignment, use a reference standard that shares provenance with the retrieval corpus and run on a release that preceded the deployed gate. We therefore pre-register OncoSEAL (Sealed Evaluation of Agents in the Loop), combining a retrospective set of over 1,000 cases sealed by time, a within-period concurrent control that splits by reader, an analysis plan hashed at the release tag and an external cohort with an independent reference. The seal is structural, since deployment experience can re-enter the system only through an interface unavailable on sealed cases by construction. This paper precedes the first scoring, so every number in it is a development-set or harness-contract measurement, and we make no claim about the agent’s accuracy in deployment.