In search of a definition of importance: Do attributions capture it?
Abstract
Attribution methods are extensively used to identify which parts of the input are important for a model's decision. However, despite recent efforts there is still no consensus on what exactly they capture and how to evaluate them. We start addressing this problem by explicitly modelling the implicit assumptions that sit at the foundation of attribution methods. We define two types of importance contributions an input region can have. We then create a framework that decouples the process of verifiably establishing contributions from that of evaluating attributions. This shows that current attribution methods struggle to reliably capture input contributions. Further investigation shows a stark difference in attribution performance between model behaviours, exposing a previously overlooked aspect in the attribution literature. Our work calls for clear, quantifiable statements about what attribution methods aim to capture, along with rigorous evaluation frameworks.