Phylogeny-weighted conformal prediction for genomic antibiotic susceptibility testing: four failures of a natural idea
Abstract
Split conformal prediction gives distribution-free coverage under exchangeability. Bacterial isolates are not exchangeable - they are related by a phylogeny - so when a novel lineage reaches the clinic, the coverage guarantee on a genomic antibiotic- susceptibility model should be expected to fail. It does: over 491 (drug, clade) holdouts on 10,227 Mycobacterium tuberculosis genomes, realised coverage falls below 0.70 for 6.9% of pairs against a 0.90 target, and a model's AUROC on a held-out clade carries essentially no information about whether its guarantee survived (within-class Spearman rho = -0.005). The natural fix is to reweight or gate conformal prediction by phylogenetic distance, which is published, cheap and available. We report that this does not work, in four independent ways: (i) normalised tree-kernel weights are provably and empirically degenerate under deep clade shift, because the test point's distance factors out; (ii) the shift is generalised rather than label- or covariate-only, so no reweighting scheme helps - including one handed the true test prior; (iii) phylogenetic distance used as an abstention score is anti-correlated with error, performing worse than random abstention; and (iv) phylogenetic novelty does not predict coverage failure, with every phylogeny- and novelty-derived feature below 0.02 gradient-boosting importance. What does work is unrelated to phylogeny: the model's own predicted-probability distribution on the unlabelled batch. We also show that in the binary setting conformal abstention is exactly a confidence threshold, so conformal prediction contributes no ranking information at all - only a calibrated threshold, which is precisely what clade shift destabilises.