SCAPE: Single-Cell Analysis Proficiency Exam
Abstract
Large language model agents can execute increasingly complete single-cell and single-nucleus RNA-sequencing workflows, but their evaluation must distinguish scientific quality from workflow completion and run-to-run stability. Existing benchmarks commonly grade isolated endpoints, prescribed reference artifacts, or expert-authored process rubrics. We introduce SCAPE, a hierarchical evaluation framework that propagates data-derived and reference-concordance evidence through four nested resolutions: oriented metric utilities, analysis-step scores, processing and analysis scores, and an overall quality score. Metrics computed from the input and agent-produced artifacts provide dense local signal, while selected outputs are compared with curated results from peer-reviewed analyses. Repeated equivalent runs separately quantify step and end-to-end reliability, quality dispersion, and agreement between scientific outputs. This hybrid design permits multiple defensible analysis paths, localizes failure modes, and reduces dependence on dataset-specific answer labels while retaining an endpoint-level check on the resulting biological conclusions. Across 150 executions on three datasets, reliability ranged from 33.3\% to 100\%; dataset-balanced quality differed modestly, while PBMC rerun variation was comparable to between-model variation.