Beyond Incremental Beam Search for Compositional Explanations of Neurons
Abstract
Compositional explanations describe the alignment between neuron activations and concepts through logical formulas, obtained by solving a search problem over combinations of concepts. Due to the combinatorial nature of this space, prior work relies on incremental beam search, which restricts the exploration but limits the expressiveness of explanations and may discard intermediate candidates needed to identify higher alignment. In this paper, we propose a non-incremental beam search framework that expands the search space while keeping the computation tractable. Our approach introduces a non-incremental beam expansion strategy and a functional-equivalence-aware beam search that leverages equivalence classes of formulas to improve both search quality and efficiency. Experiments on natural language and vision tasks show that our method finds explanations expressing a higher alignment than the incremental formulation while maintaining comparable runtime, and leads to a significant increase in the number of optimal explanations when these can be computed. Overall, these results show that expanding search beyond incremental formulations leads to more expressive and accurate explanations, making our method a promising replacement for standard incremental beam search.