Espalier: Sparse Causal Interfaces from Joint First-Order Geometry
August Willoughby
Abstract
Circuit discovery is conventionally evaluated by whether an ablated subgraph reproduces a behaviour on its own. Yet the machinery used to find circuits, including activation patching and its gradient-based approximations, measures something else: whether a retained set transmits a specified counterfactual change. We formalise the second criterion promptwise as a finite direct-mediation defect and find that the two are not interchangeable: across prospectively frozen fixed-cardinality interface families on two GPT-2 Small tasks they reorder 14\% to 42\% of interface pairs; on IOI, changing only the mean-ablation reference materially changes the amount of disagreement. Taking the second criterion as the target, we ask what selection should preserve. Existing methods reduce each candidate to a scalar attribution and rank candidates independently, discarding how their effects combine across the perturbation distribution. Espalier instead retains a promptwise contribution vector per candidate and scores whole interfaces by an interface defect, the intervention-scale-free first-order limit of the finite estimand. The surrogate ranks held-out causal error within frozen fixed-cardinality interface families at $\rho \geq 0.988$; and holding the contribution matrix and search procedure fixed, retaining cross-candidate terms lowers held-out defect at $8/8$ IOI and $8/9$ Acronym budgets. At matched head budgets, Espalier attains lower held-out defect than EAP, EAP-IG and AtP* in all tested cells, and lower than AtP in 13/14 with one identical-selection tie, with the largest gaps in the sparse regime. Low defect does not confer standalone sufficiency: under mean ablation Espalier's IOI interface ranks last of the five selectors. Which criterion a discovery method should optimise, and which reference distribution operationalises it, is an empirical commitment rather than a definitional one.
Chat is not available.
Successful Page Load