Anytime-valid inference for calibration of posterior approximations
Abstract
Simulation-based calibration (SBC) checks are standard practice for assessing posterior approximations. While they are useful diagnostics for under- or overdispersion, interpreting them as calibration guarantees has two issues. First, the direction of inference does not match the question of interest: a lack of evidence against the null of exact calibration cannot be interpreted as evidence for it. Second, an approximation trained via optimization is typically not exactly calibrated, so any consistent test rejects with probability tending to one as the simulation size grows. To solve both problems, I propose a test of the reversed null that the calibration error exceeds a tolerance. The test is based on an e-process construction and stays valid under optional stopping, so the researcher can keep simulating to accumulate evidence. It accommodates arbitrarily many test quantities without multiplicity correction.