Conclusions from Circuit Extraction Depend on the Level of Description: A Controlled Comparison
yang sheng ⋅ Jie Fu
Abstract
Circuit extraction identifies model components that preserve a target behavior under ablation, yet it remains unclear which parts of the reported circuit—components, edges, or coarser summaries—remain stable across reasonable extraction choices. We study this question in a Lean-style, rule-generated tactic-prediction benchmark spanning atomic and compositional proof tasks. The proof-state rules and task structure are fixed by construction, and dense and weight-sparse variants of one transformer architecture are compared. In this controlled setting, we compare three ways to describe a circuit: a compact prediction-preserving circuit, a broader graph that keeps read, write, and routing components around that circuit, and a pruning graph whose size is set by a target post-ablation loss budget. We vary model sparsity, whether the query and key sides of attention are represented together or separately, the post-ablation loss budget for the pruning graph, and which supervised checkpoint initializes reinforcement learning (RL). When the same extraction method is used, dense and weight-sparse checkpoints overlap far more in the set of attention heads identified than in exact component-to-component edge lists. This head-level overlap is well above a random top-$k$ baseline; the exact-edge overlap is not consistently so. Tracking query and key sides separately, and varying the loss budget across the tested range, preserves the ordering of the RL initialization conditions when each graph is summarized by how many components it selects. Among these RL conditions, the largest gains on compositional tasks are accompanied by the largest fraction of compositional-task circuit nodes that lie outside the matched atomic-task circuits. These findings show that, even in this controlled setting, behavior preservation under ablation does not determine a unique circuit-level conclusion: the conclusion depends on which graph is reported, how it is extracted, and at which level of description the comparison is made.
Chat is not available.
Successful Page Load