Contaminated Ancestors: Seven Ways Our Own Evaluation Protocol Nearly Told Us the Wrong Thing
Christopher Y Huang ⋅ Ryan J Ahn ⋅ Wenhao Lu
Abstract
We report a full-length audit of a single scientific-ML evaluation protocol, conducted on our own work. The task is detecting chimeric ancestral protein reconstructions: given a family whose history contains recombination, does a scoring model plus a permutation test flag the contaminated segment? The headline metric was detection rate. We show that on this problem detection rate has essentially no construct validity: at $n = 15$ conditions the test model fires $2/15$ and the null control fires $2/15$, because a correctly calibrated permutation test is supposed to fire near $\alpha$. The measurement that does discriminate is localization (mean Jaccard $0.81$ vs $0.08$ among firings), and the two dissociate in at least three independent places in our data. We then document six further ways the protocol produced or nearly produced a wrong conclusion: a sign convention that inverted a metric (AUC $0.048$ where the truth was $0.952$); a metric whose null is $0.50$ to $0.68$ rather than $0.5$, so that $0.5$ is the wrong baseline; a two-to-three seed regime in which seed variance exceeded every treatment effect and a $0/6$ control looked immaculate before a $2/12$ rerun showed it firing at nominal rate; a confounded comparison in which the treatment covaried with the number of input sequences; a competitor benchmarked at an undisclosed structural disadvantage; and an internal summary artifact reporting a run whose data was never committed, which could not be regenerated because the aggregation script had silently drifted from the result schema. Separately, we recorded an accelerator backend corrupting tensors mid-run without raising, converting $p = 0.030$ into $p = 1.000$ under identical settings. We close with a transferable protocol checklist. Every failure here is generic; none required the biology.
Chat is not available.
Successful Page Load