Benchmarking Single-Cell Decoders: Rankings Depend on Metric and Readout
Abstract
Evaluations of single-cell expression reconstruction combine paired cell-level accuracy, gene-level distributions, and population heterogeneity. Although these tasks ask different questions of decoders, reconstruction benchmarks often report a single model ranking. We show this ranking is not a fixed property of the decoder and that it depends jointly on what the decoder predicts, what its latent space assumes, and how its output is read out at evaluation time. We disentangle these by extending ReconEval, a recent reconstruction benchmark, with a broader metric suite. We also introduce a hurdle-continuous ranked probability score (CRPS) decoder, which combines a deterministic latent representation with a distributional reconstruction. Across four single-cell datasets, point-estimate decoders score higher on cell-level accuracy, while distributional decoders better preserve gene-level distributions and population heterogeneity. This ranking also depends on how the output is read. For the same trained model, sampling from its predicted distribution performs better at the gene level, while taking the mean of that distribution performs better at the cell level. Reconstruction performance thus depends on the metric, biological scale, and readout used for evaluation.