Explaining Cross-Modal Model Behavior with Gradient-Estimation-Based Feature Interaction
Abstract
Multi-modal machine learning has become a dominant paradigm for building large-scale models that integrate information across heterogeneous modalities and achieve unprecedented representational capabilities. However, the multi-modal nature of such models introduces unique challenges for explainability. Specifically, the standard feature attribution task considers only individual feature contributions. While solutions of this kind can be adapted to operate on each modality, such first-order explanations are insufficient to characterize cross-modal feature interactions, which are essential for understanding model behavior involving cross-modal alignment. To bridge the gap, this paper studies higher-order feature interactions in multi-modal settings and proposes Gradient-Estimation-Based Feature Interaction (GEFI). GEFI is derived from the proxy gradient estimation framework, and we establish its theoretical equivalence to the Shapley Interaction Index at arbitrary orders. Beyond the general theory, we instantiate GEFI-PM for CLIP-like models by exploiting their dual-encoder architecture. The resulting formulation enables the efficient estimation of interactions between visual and textual features. Compared to existing methods for revealing cross-modal interactions, which generally operate on fixed image patches, GEFI offers fine-grained pixel-level interactions while maintaining computational efficiency, yielding more expressive results.