Evaluating Predictive Uncertainty for Mature-Oocyte Counts in IVF
Abstract
Two predictive distributions can have the same expected mature-oocyte count yet assign substantially different probabilities to poor maturation. We tested this in 2,093 retrieval cycles from 1,339 Brigham and Women's Hospital (BWH) patients, including conventional-IVF, ICSI, and split-insemination cycles. We combined one fitted mean prediction with two count distributions. The standard binomial imposed its usual count variance; the beta-binomial allowed additional cycle-level variation. Among cycles with at least ten retrieved oocytes, 15.0% had no more than half of those oocytes reach metaphase II. The standard binomial model predicted this poor-maturation outcome in only 6.5% of cycles, whereas the beta-binomial predicted 17.0%. Thus, despite predicting the same average number of mature oocytes, the standard model underestimated the risk of a poor cycle by more than twofold. The discrepancy remained after restricting BWH to cycles that produced a freezable blastocyst and appeared in a separate success-selected cohort. Because mature-oocyte yield constrains subsequent fertilization and embryo yield, understated tail risk can propagate into falsely confident downstream forecasts. More generally, we show that biological ML systems can fit an average while misrepresenting variation among outcomes nested within the same patient or sample.