Calibration Selection, Not Contamination, Breaks Conformal FDR Control in Graph Anomaly Detection
Abstract
We study a failure mode of conformal false discovery rate (FDR) control for graph anomaly detection in which the guarantee is violated without any visible sign in the procedure's output. The standard precaution against calibration contamination, restricting calibration to normal nodes with no anomalous neighbors, is itself a covariate filter. On a large fraud-detection graph, this filter shifts calibration toward a far lower degree than the population under test. When a detector's score depends on degree, the guarantee breaks: realized false discovery rate reaches 0.785 against a nominal target of 0.10, and the resulting discovery set is indistinguishable in its output from a valid one of the same size. Holding the graph fixed and varying only the detector shows that the failure tracks a property of the detector, not the graph, replicating on a second graph and disappearing on a third where the underlying covariate relationship does not hold. The threat this filter is meant to prevent, actual contamination, produces no comparable failure when we test it directly. Inverse-propensity reweighting, the standard correction for a known covariate shift, does not reliably restore validity either, and worsens two of the five configurations where it has anything to correct. A label-free score gap tracks the violation closely enough to serve as a deployment-time check, and a training-free degree lookup matches most trained detectors across four benchmark graphs, raising a direct question about what these benchmark scores have been measuring. These results indicate that a defense against one failure mode does not provide evidence against another, and that this distinction can be checked using summary statistics that are already available at deployment.