Finite Sample Power Analysis for Selective Medical AI Audits
Abstract
A high AUROC does not ensure that a clinical audit can establish both low false omission risk and sufficient coverage for a selective cancer triage policy. We therefore define audit readiness as the probability that a future audit will provide enough evidence that both requirements are met. This probability depends on how often a policy clears malignant and benign cases—information that AUROC and sample size alone do not provide. We estimate these clearance rates from a separate labeled sample to calculate readiness when the score distribution is unfamiliar. For audits requiring 95% confidence that both requirements are met, Empirical CDF readiness estimates had a mean absolute error of 2.11 percentage points across individual simulated assessments, compared with 5.97 for an AUROC informed binormal model and 7.85 for a Fitted binormal model. Across three clinical datasets, analysis of fixed prediction scores identified one design that reached the readiness target at a practical audit size, another requiring much larger samples, and a third in which no threshold met both population requirements. Audit readiness identifies when a larger audit can support a clinical claim and when the model or policy must change.