Measure the Baseline Before You Beat It: A Reproducible, Confidence-Interval-Grounded Suite of Classical Quantum-Error-Correction Decoders
Abstract
Claims that a learned decoder “beats” a classical quantum-error-correction (QEC) decoder are only as trustworthy as the classical bar they are measured against—and that bar is often reported without confidence intervals, without an online-latency distribution, and with logical-error-rate (LER) points compared across noise models that are not comparable. We release a reproducible baseline suite that fixes these three gaps. It evaluates the standard classical references—minimum-weight perfect matching (MWPM) for the surface code and belief-propagation with ordered-statistics decoding (BP-OSD) for bivariate-bicycle (BB) qLDPC codes— across three explicitly non-comparable arms (circuit-level surface, code-capacity BB, phenomenological space-time BB). Every LER carries an exact Clopper–Pearson binomial 95% confidence interval and its failure count, and under-resolved rare-event points are reported as bounds, not values. We add online, single-shot decode latency (p50/p95/p99), which is distinct from batch throughput and is the operational figure a real-time decoder must meet; ordered-statistics fallback gives the phenomenological decoder a heavy right tail (p99/p50 ≈ 26). Finally, we include honest negative controls: off-the-shelf neural sequence models under a fixed compute budget fail to decode these codes by one to four orders of magnitude, and we scope what a genuinely competitive learned decoder would require. We make no competitive-decoder claim; the contribution is a rigorously-measured, seed-frozen bar so that future learned-decoder papers can be evaluated fairly, per arm, with uncertainty