Learning Global Probabilistic Explanations
Frederic Koriche ⋅ Louenas Bounia
Abstract
Interpreting the predictions of complex black-box classifiers remains a central challenge in explainable artificial intelligence. While local explanations clarify individual predictions, there is a significant need for global probabilistic explanations that capture a model's overall behavior across the input distribution. We define such an explanation as a small subset $K$ of features, whose relevance is measured by the probability that the classifier assigns identical labels to two independently sampled inputs that agree on $K$. Based on this notion, we investigate the task of identifying maximally relevant global explanations under a cardinality constraint, focusing on product distributions with full support on discrete feature spaces. We introduce a spectral explainer that leverages membership queries to the black-box classifier and employs Fourier-analytic, junta-learning methods to produce probably approximately correct (PAC) explanations. Our algorithm is fixed-parameter tractable with respect to the explanation size limit, maximum feature cardinality, and desired accuracy. Experimental evaluations on synthetic and real-world datasets show that the spectral explainer provides highly relevant global explanations and outperforms heuristic greedy methods as the explanation size increases.
Chat is not available.
Successful Page Load