Eight Donors, Not a Thousand Cells: Auditing Single-Cell Analysis Agents for Pseudoreplication
Abstract
Autonomous agents now analyze single-cell data and report differential-expression "discoveries". They are graded on their answers, not on whether their methods are valid. The best-known failure in this setting is pseudoreplication: treating cells from the same donor as independent replicates. Recent audits of analysis agents build null datasets by shuffling labels across rows. These shuffled nulls cannot detect pseudoreplication. We treat an analysis agent as a statistical procedure and audit it with donor-level nulls. We build exact nulls by relabeling whole donors of real control data. We plant known effects with binomial thinning. We also introduce permutation replay, which re-runs the agent's own script under donor relabelings. We ran a pre-registered study on real PBMC data on a single laptop. Shuffled nulls made a cell-level Wilcoxon test look calibrated (family-wise error rate, FWER, 0.00). Donor-level nulls showed that the same test made false discoveries on every draw (FWER 1.00). A local coding agent (Qwen3-14B) reported false discoveries in every valid run on donor-level nulls. The median was 75.5 false genes per run. The agent always tested cells instead of donors. Telling the agent that the donor is the unit of replication changed nothing. A typed pseudobulk tool cut false discoveries to a median of 0, but FWER stayed at 0.38. On null data, replay removed every false discovery. With only four donors per group, it also removed nearly all planted effects.