PerturbReason: A Knowledge-Grounded Benchmark and Framework for Cell-State–Conditioned Mechanistic Reasoning of Perturbation Effects
Abstract
Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PerturbReason, a knowledge‑grounded benchmark for cell‑state--conditioned reasoning about perturbation effects. It tests the models to generate mechanistically faithful explanations and assesses their robustness against complex shifts, such as new cells, unseen perturbations, and cross-modal or combinatorial extrapolations. PerturbReason combines single‑cell genetic and chemical perturbation data across multiple cell lines with knowledge graphs, and dynamically conditions pathways on cell‑specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between predictive accuracy and mechanistic reasoning. Specifically, these models exhibit failure modes largely invisible to standard benchmarks, such as deriving correct answers through flawed logic, ignoring cellular context, and generating directionally inconsistent mechanisms. As a reference probe of the benchmark, we present PerturbRM, a large language model trained to align outcome predictions with context-specific regulatory reasoning. PerturbReason thus provides a rigorous diagnostic benchmark for studying and improving faithful, generalizable reasoning in data-rich scientific systems.