Certified Robust Interpretability via Concept-Space Stability under Interventional Proxies
Yanming Cheng ⋅ Di Zhang
Abstract
A concept-bottleneck classifier can remain correct, remain $\ell_2$-robust in pixel space, and still silently rewire which concepts carry its decision under a small perturbation. The label looks safe; the audit trail has been swapped out underneath it. This failure mode is not a quirk of randomized smoothing: any certifier that sees only the input and the label distribution admits classifiers with arbitrarily large concept drift inside its certified ball. To certify the missing object, we propose CRISP, which delivers a deterministic concept-space certified radius $r^\star(x)$ inside which both the label and the concept activations are stable. The radius is given by a closed-form Lipschitz--margin bound and is tight up to a factor of two on the relevant encoder/head/margin triple. Under $\varepsilon$-concept-faithfulness and a known concept-to-target subgraph, $r^\star$ further certifies the causally meaningful concepts; a matching impossibility result shows that faithfulness cannot be removed. Training relies on interventional proxies---SCM simulators, attribute editors, environment swaps, concept-matched retrieval---which approximate $\mathrm{do}(\cdot)$-interventions on natural images. We bound the proxy-vs-truth gap in closed form and measure it directly in the synthetic regime where ground truth is available. Empirically, the bound is informative rather than vacuous. On a synthetic SCM with observable $c^\star$, 923 correctly-classified samples (5 seeds, $k \in \{3,6\}$) yield zero certificate violations under budget-5.0 APGD, with empirical tightness $r^\star/r_{\mathrm{adv}} \approx 0.32$. Concept-fidelity (MCC) rises from $[0.37,0.51]$ for plain CBM to $[0.74,0.83]$ for CRISP, and concept-space attack success drops by up to 13 points at matched budget. A frozen-backbone Waterbirds run preserves the zero-violation property across 1,980 samples and surfaces what we call the Lipschitz tax: a $\approx 27\times$ gap between architectural and local encoder Lipschitz constants that is the dominant bottleneck to tightness on natural images. A preregistered protocol for CUB-200 and CANDLE (Reddy et al., 2022) is fixed in Appendix C.8.
Chat is not available.
Successful Page Load