Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery
Abstract
Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly interpreted as evidence of distinct mechanisms. We test this assumption by drawing input tokens from bands defined by their frequency in the pretraining data while holding the task fixed. The discovered circuits appear specialized by frequency when compared structurally, but functional and representational analyses show no reliable evidence of corresponding differences in their computations. We term this mismatch phantom specialization. Using the Literal Sequence Copying task across four frequency bands plus a control sampled according to token frequency, we extract 75 circuits from five Pythia models (70M-1.4B). We find that structurally distinct circuits implement the same computation: band-specific edges transfer broadly across bands, a core shared across most bands recovers at least 99% of circuit performance in models above 70M, and causal interchange interventions confirm that internal representations are interchangeable across frequency bands. A smaller replication on a subject-verb agreement task shows the same pattern: circuits differ structurally, transfer broadly across bands, and the core shared by most bands recovers nearly all of the circuit's accuracy in all five models. Repeated extractions within the same frequency band further suggest that discovery algorithms sample from an equivalence class of valid subgraphs rather than recovering a unique mechanism. Standard evaluation practice obscures this pattern: source-level evaluation inflates apparent faithfulness, while edge-level evaluation reveals the many-to-one mapping from structure to function. We find no reliable evidence that frequency bands are processed by different computations. In this work, we do not test specialization at the level of token positions or of features inside a component. Our results show that structural differences between circuits are not sufficient evidence for distinct mechanisms, and that exposing this requires edge-level evaluation and cross-condition transfer tests.