Improving Causal Explanations
Abstract
Explaining model predictions in high-stakes settings requires knowing why a prediction was made, a central challenge in modern AI. Yet existing methods lack a principled causal framework for explanation, leading to a disagreement problem: explanation methods produce conflicting attributions with no way to determine which is correct. Based on the literature in cognitive science, law, and philosophy, and using the language of causality, we formalize four desiderata for explanation methods that assign numerical credit to variables. We show that no prior explanation method including SHAP and LIME satisfies all desiderata, and prove the Explanatory Impossibility Theorem, which shows that even in principle, no method can satisfy all desiderata without causal assumptions. We introduce counterfactual Shapley values, an aggregation of a novel counterfactual quantity we call the natural total effect, and prove they satisfy all desiderata. We develop an algorithm for computing tight bounds on counterfactual Shapley values. Experiments corroborate our theory.