Worse Than Random: J-Space Ablation Is Anti-Selective at Every Measured Qwen3 Scale
Peter Flo
Abstract
The global-workspace interpretation of J-space rests on a selectivity claim: ablating a model's "verbalizable" directions should impair internal reasoning while sparing routine language. Across a historical 4B pilot and three preregistrations on the open Qwen3 family, we find the opposite signature at every measured scale. J-ablation damages ordinary-text prediction **more** than a norm-matched random intervention on 50 of 50, 49 of 50, and 42 of 50 paired WikiText sequences at 1.7B, 14B, and 32B (exact sign tests, $p\le10^{-6}$). A selectivity index, the ratio of J to matched-random damage, declines monotonically across the four scales, from 3.38 at 1.7B to 1.16 at 32B, yet never reaches parity; the 1.7B and 32B endpoints are confirmatory under separate preregistrations. The recorded displacement norms explain why the control fails: matched perturbation norm at the injection site does not deliver matched severity. The J-perturbation attenuates downstream while causing about $2.4\times$ more damage per unit displacement, so norm-matched random controls are not severity-matched controls. A preregistered conjunctive validity gate stops at every scale, and two of four stock lenses fail its lens prerequisite. Because no finished stock 32B lens is available, we fit one from scratch with the source study's released recipe; the resulting lens, the best-validating in the study, still yields anti-selectivity. The results do not test the source study's frontier-scale models: the attenuating index is consistent with selectivity emerging only beyond 32B, a falsifiable boundary for future instruments.
Chat is not available.
Successful Page Load