What Graph Metrics Actually Measure in Feature Co-activation Graphs
Samanyu Badam
Abstract
Graph measures such as treewidth, separators, modularity, and clique structure might track how a model's computation is organized, and could then support detectors without labeled examples of every behavior of interest. We test this idea on a case where it first appears to work and then fails under pre-specified controls. On GPT-2 small with a residual-stream sparse autoencoder, 18 graph measures separate factual prompts from prompts eliciting confabulation or sycophancy (random-forest AUC $0.851$; five-seed mean $0.832$). However, terminal punctuation alone classifies every prompt correctly, and the graph measures largely track one quantity, active-set overlap, which together with prompt length reaches AUC $0.829$. The graph construction also limits what some measures can mean: each co-activation graph is a union of cliques over $\kappa$-subsets of positions, the graph has no source or sink for interpreting vertex separators, and treewidth closely follows clique number in these data. We turn these problems into four validation checks: simple input-level baselines, matched prompt pairs, threshold sweeps, and simulated null graphs. Under an internal pre-registration, the matched pairs erase every effect and none survives correction (a probe control limits what that null can show). The threshold sweep shows that the edge-count effect peaks at the study's chosen threshold. A frequency-preserving shuffle shows that per-feature base rates explain most of the failure of the registered null model, reducing edge $z$ by an order of magnitude while leaving a smaller genuine component of within-position co-activation. On every contrast we test, a rule based only on the input matches or outperforms the graphs. Applied to an external verifier's released data, the one reported result whose labels we can reconstruct retains margins of about $3$ AUROC and $13$ FPR@95 points over pre-specified text-and-position rules. The other label set is ambiguous: the pre-specified rule matches the reported point estimate within CI under one reading and falls short under the other, while post-hoc variants exceed it under both.
Chat is not available.
Successful Page Load