BrainEM: A Large-Scale and Diverse Benchmark for EM Neuron Segmentation in Connectomics
Abstract
Connectomics requires dense segmentation of large-scale electron microscopy (EM) volumes using models trained on small labeled blocks. In practice, a connectomics lab facing a new raw volume has two options: annotate a small block and train on it, or skip annotation and train on the union of available public labeled datasets. We introduce BrainEM, a benchmark that formalizes these options as two standard tracks: the Same-Source Track (Annotate-and-Train) and the Cross-Source Track (Reuse-and-Train). BrainEM brings together 16 EM datasets, including one in-house mouse dataset that we contribute, spanning multiple species, imaging modalities, and resolutions. We further add large-scale test volumes that bring the evaluation closer to realistic practice, and four diverse test targets in the Cross-Source Track. The data, post-processing, and evaluation metrics are fixed across all methods, and the framework is fully model-data decoupled so that new methods can be added with minimal effort. We conduct a systematic evaluation of six representative baselines covering the main technical directions in EM neuron segmentation, making BrainEM the first systematic evaluation at large scale and across diverse cross-source targets. The results differ from prior small-scale evaluations: rankings on the Same-Source Track no longer match those reported on small CREMI splits, and methods diverge sharply on the Cross-Source Track. Both tracks still leave clear room for improvement. We release BrainEM as a public platform to support unified comparison and faster method development in connectomics. All datasets and code are available at https://github.com/kwinderic/BrainEM