Decision Path Tracing for Causal Analysis in Transformers
Abstract
Investigating the causal structures of transformers often entails a trade-off between the granularity of the explanation and computational overhead. Existing approaches either bypass the full internal chain, potentially missing critical decision-relevant links, or struggle with the combinatorial explosion of path enumeration. To address these challenges, we propose an efficient automated path-tracing method that ensures both causal reliability and algorithmic feasibility. Through extensive verification, we find that (i) the identified paths are the primary sources for decision signals, as evidenced by a marked decrease in self-repair compared to non-path components; and (ii) transformers utilize modular, class-shared path mechanisms to process categorical information. Beyond these findings, we show that the potential implications of our method can stem from sparse manipulation, highlighting efficient applicability to downstream tasks such as model debugging and pruning.