Marginal Conformal Coverage Hides Per-User-State Gaps in Return-Time Forecasting
Abstract
We study conformal uncertainty quantification for user return-time prediction. Conformal intervals provide marginal coverage guarantees, but these guarantees need not hold conditionally on user state. On 2.45M listening sessions from Yambda-50M, with a 30-day label follow-up and calibration on held-out users, intervals calibrated to 80\% marginal coverage cover only 67\% of sessions from users with declining activity and 66--70\% of sessions from users with fewer than five previous sessions. These gaps persist despite stable population-level behaviour throughout the untruncated weekly evaluation period, indicating substantial heterogeneity in conditional coverage across user states rather than a simple temporal-drift effect. Two models that explicitly condition on the user's timing history (a gradient-boosted quantile regressor with hand-crafted return-time features and a Transformer density model over gap tokens) raise worst-cohort coverage to approximately 0.78, with interval widths comparable to the strongest heuristic baseline and lower prediction error. Mondrian calibration over predefined activity and history cohorts improves weaker predictors, but residual failures remain in sparse calibration cells and in cohorts outside the chosen partition. We find no consistent gain from item information: learned item embeddings and joint next-item training degrade conditional coverage, while frozen content embeddings perform similarly to the gap-only model. The main findings replicate on Yambda-500M and, under a reduced protocol, on KuaiRand-27K, where an additional calendar shift also affects marginal coverage.