Calibrated Where the Answer Was Already Known: Conformal Fate Sets Miss Exactly the Cells Whose Fate Is Undecided
Abstract
We set out to give cell-fate prediction a distribution-free reliability layer: wrap a fate predictor in split conformal prediction and hand a biologist a set of fates that provably contains the truth 90 percent of the time. On lineage-traced hematopoiesis (LARRY, 1,408 clones) it works exactly as advertised, and it is close to useless. Marginal coverage lands at 0.907 against a nominal 0.90. But coverage falls from 0.966 on clones whose fate was already settled by day 2 to 0.858 on clones that were still undecided, so the guarantee is paid for by the cells nobody needed a prediction set for and it fails on the cells the experiment was run to study. The textbook repair does not work either: whether a clone's fate is decided is nearly invisible in the day-2 transcriptome, at AUC 0.594, and group-conditional calibration on that predicted group moves the undecided stratum only from 0.854 to 0.867. The diagnosis is not that conformal prediction is the wrong tool, and the same data shows why. The identical features, model and split predict which fate a clone takes at a mean AUC of 0.887, and calibrating on that predicted majority fate does repair its coverage deficit (basophil clones, 0.84 to 0.88). Direction is legible in this assay and determinacy is not. The toolkit repairs what the measurement can see and cannot repair what it cannot, so any wrapper conditioned on the same day-2 state inherits the same ceiling.