DD-CAM: Minimal Sufficient Explanations for Vision Models Using Delta Debugging
Abstract
Class Activation Mapping (CAM) and related saliency techniques are widely used to explain predictions of deep vision models, producing post-hoc heatmaps that highlight image regions a model relied on. However, existing CAM methods aggregate weighted contributions over the full set of internal units, including units whose contribution may be incidental. The resulting heatmaps are often diffuse and obscure which units are actually sufficient to preserve the prediction. We reframe visual explanation as the selection of a 1-minimal sufficient subset of internal representational units, feature maps in CNNs and patch tokens in ViTs, whose joint activation preserves the model's prediction. A subset is 1-minimal sufficient if removing any single unit changes the prediction. We propose \textbf{DD-CAM}, a gradient-free procedure adapted from delta debugging in software engineering that instantiates this formulation. DD-CAM exploits classifier-head structure with two regimes: iterated single-unit removal for non-interacting heads (GAP+FC), and recursive partition-and-reduce for interacting heads (multi-layer FC stacks and ViT self-attention). On 2{,}000 ImageNet validation images across eight CNN and ViT architectures, DD-CAM is the strongest method on 13 of 18 faithfulness metric-group cells against thirteen baselines. On the NIH ChestX-ray14 bounding-box set, DD-CAM improves IoU by 45\% and F1 by 22\% over the strongest baseline while producing single-region explanations. These results suggest that necessity-grounded internal-unit selection can produce more focused and faithful explanations than weighted aggregation over all units, with particular promise for safety-critical settings such as medical imaging.