scShapeBench: Discovering geometry from high dimensional scRNAseq data
Abstract
High-dimensional point cloud data arise across many scientific domains, notably single-cell biology. The "shapes" or topologies of these datasets are informative of types of information that can be extracted from the datasets. For example clustered data admits the extraction of cell types or cell states in a static analysis of the datasets. Continuous trajectory structures admit continuous transition or trajectory analysis, while other shapes such as archetypal shapes admit continuum extraction with a range of cells spanning behaviors. While analysis pipelines exist, they often presuppose shape in data. For example, the standard Seurat pipeline combines UMAP visualization with Louvain clustering. This assumes clustered data. Tools like Monocle and Spade assume a tree-like shape, and flow-models like MIOFlow and Conditional Flow Matching are suitable for trajectories. Deciding which pipeline to apply to which part of the data is often the realm of bioinformaticians who visualize and qualitatively analyze the data before selecting one. However, with the advent of agentic AI scientists, it becomes important to automate data shape detection, particularly into categories that are relevant for downstream analysis pipelines. Towards this end we introduce \textsc{scShapeBench} a benchmark dataset comprising both synthetic and single-cell expert-annotated datasets that are meant for the task of shape detection. Synthetic datasets are sampled from a ground truth “skeleton graph” with variance. Real single-cell datasets are curated from a variety of sources and are annotated by experts, classifying four categories; clusters, single trajectory, multi-branches and archetypes. In addition, we provide a baseline method, scReebTower, to bridge the gap between data visualization and pipeline selection. scReebTower relies on the diffusion geometry to extract Reeb graphs. We provide new topology-aware metrics with which we evaluate scReebTower and existing methods PAGA and Mapper on synthetic data. On single-cell data we curate expert annotations of shapes, and showcase evaluations of methods. Our comparisons indicate scReebTower outperforming other baselines. Overall, our contributions span benchmarks, evaluation metrics, and a novel baseline method for automated shape detection in high-dimensional single-cell data.