Causal Concept Explanations for Deep Neural Models
Abstract
Concept-based explanations in neural models tie their outputs to human-meaningful concepts that users can understand, troubleshoot, and trust. But these explanations are typically correlational rather than causal: a concept can fire alongside an output without driving it, yielding interpretability that looks right but misleads. We introduce Causal Concept Wrapper Network (CCW-Net), the first use of mediation analysis as a training objective in deep networks. We further apply mediation analysis as an evaluation tool to quantify, per sample, how much of any model's output is causally attributable to each concept. As a training objective, CCW-Net shapes a model's concept embeddings such that each concept's contribution to the output is both locally necessary and sufficient for the share it is credited with, and independent of contributions from other concepts. Counterfactual concept samples drawn from a learned per-concept flow anchor a set of causal alignment criteria: necessity drives the effect of every irrelevant concept toward zero; sufficiency requires relevant concepts to fully reconstruct the model's output; and independence enforces that the total effect decomposes linearly across concepts. We evaluate CCW-Net across aircraft control, driving, and fine-grained image classification. The result is a step toward concept-based explanations that are not merely interpretable, but quantifiably causal, supporting safer deployment, evaluation, and certification of transparent and trustworthy neural systems.