BioSafetyBench: Agentic Cascade Evaluation for the Bio-AI Ecosystem
Abstract
Biological risk is often a property of a cascade rather than a single model output. A generated artifact may become consequential only after transcription, translation, folding, binding, pathway propagation, or host-directed off-target interaction. Existing bio-AI safety evaluations largely grade each model within its own modality, missing downstream risks and inheriting self-grading circularity when model and evaluator share representations. We introduce BioSafetyBench, an agentic cascade-evaluation platform that treats the bio-AI ecosystem as the unit of safety evaluation. BioSafetyBench routes each model-output pair to one of 15 standardized probes over a six-level hierarchy from genome to organism-level outcome. Fixed bridge agents propagate outputs through two causal pathways: Pipeline A follows forward biological cascades, while Pipeline B evaluates CRISPR and siRNA designs through genome- and transcriptome-wide off-target alignment. At every reached level, external frozen evaluators compute domain-native biomarkers and compare them with literature-anchored thresholds, yielding a Cascade Concern Profile and Cascade Depth. Across 25 foundation models, BioSafetyBench surfaces risks invisible to single-modality grading: Influenza neuraminidase designs from ProteinMPNN and ESM-IF1 reach CD = 4; jailbreak prompts inflate gRNA off-target hits 9.05× over baseline, push 48 of 604 gRNAs to Critical, and concentrate Critical siRNA hits on MYC; protein mask-and-fill spans a 32× cross-model gap; GPT-4o complies on all natural-language-guided protein mutation prompts at CD = 2 while Claude Sonnet 4.5 refuses all prompts at CD = 0; and a three-predictor ADMET ensemble misses 30.3% of known clinical toxics. BioSafetyBench is, to our knowledge, the first bio-AI safety benchmark to provide cascade-level risk certificates whose scoring modules are weight- and gradient-flow independent from the model under test by construction.