Pre-hoc Structural Probing: Relational Graphs Are Linearly Decodable Before Query Processing in Autoregressive Transformers
DHRUV DAWAR
Abstract
Structural probing faces a fundamental methodological challenge: without rigorous confound elimination, a linear classifier may detect surface statistical regularities rather than genuine representational structure. We present a pre-hoc structural probing framework that eliminates eleven identified confounds through iterative benchmark construction, and introduces a Double Machine Learning (DML) nested cross-fitting procedure that prevents test-fold leakage through the confound-removal back-door endemic to standard probing pipelines. Using a controlled benchmark of two-clause transfer-of-ownership queries at three structural levels (disconnected, shared-node, directed chain), we find that in our benchmark, probe accuracy saturates before the query token is processed. Hidden states extracted after the second premise clause—at position Pos2, before the QUERY token—yield Probe E accuracy of 0.964–0.977 across three model families (Mistral-7B, Qwen2.5-7B, Phi-2), while states at Pos1 (after one clause) yield near-chance accuracy of 0.060–0.066. The marginal contribution of the QUERY token is $\Delta \in [-0.010,-0.002]$, consistent with relational structure becoming decodable during premise ingestion rather than at query evaluation time. At $n=2{,}000$ queries, Probe E reaches $[0.991,1.000]$ (95% bootstrap CI) across all three architectures, with $p<0.002$ under full-dimensional permutation testing and Cohen's $d\in[12.99,14.18]$. No tested surface feature reaches statistical significance in isolation (permutation $p=0.128$ after entity masking). These results are consistent with early transformer layers encoding relational topology in a linearly separable, entity-distributed geometry, complementing circuit-level accounts of entity tracking and DAG-structured theories of multi-step reasoning. We note that decodability does not establish causal downstream use; core claims are restricted to the benchmark studied; and asymptotic accuracy beyond $n=2{,}000$ is not confirmed.
Chat is not available.
Successful Page Load