Same Proteins, Different Verdicts: Toxicity-Classifier Clearance Is Not a Portable Safety Measure
Abstract
DNA synthesis screening creates one of the few control points between a designed protein and its physical production. Deployed screens mainly search for similarity to catalogued sequences of concern. Generative methods can reduce that similarity while preserving predicted structural properties, so recent studies have also used academic toxicity classifiers to assess whether a redesign remains harmful. This practice assumes that classifier clearance is stable enough to support safety claims and comparisons among redesign methods. We test that assumption after an apparent success in our own protein redesign study failed to reproduce. We rescored the same historical sequences with a specified ToxinPred3 version and found that it labelled 94 of 100 sequences non-toxic for each of two redesign methods, removing the original separation. We then compared three evaluators on 176 unmodified proteins that CommEC flags. ToxinPred2 calls 156 toxic, whereas the published ToxinPred3 rule calls only 6 toxic. A fixed negative motif term changes 50 ToxinPred3 decisions. The disagreement also changes conclusions about redesign: OAE passes CommEC for 36 proteins, but random substitutions at the same positions pass 37 (McNemar p=1.0). Protein-language-model probes and sparse-autoencoder features do not provide an independent functional label. Toxicity-classifier clearance therefore depends on the evaluator, version, threshold, input type, and positive class. It does not constitute a general safety property of a redesign.