When Logical Information Emerges: Recoverability, Behavioral Expression, and Representation Stability in Language Models
Vinayak Raj Urs ⋅ Harika Mahapatra
Abstract
A high-accuracy probe establishes that a target is recoverable from a representation, but not that the target is expressed in behavior, transfers across equivalent inputs, or is causally controllable. We study these distinctions in controlled propositional entailment using a frozen benchmark of 1,200 symbolically verified instances organized into 600 matched True/False pairs and three logical templates. Across Gemma-2-2B-IT, Phi-2, Qwen2.5-3B-Instruct, and Mistral-7B-Instruct-v0.3, nested grouped probes recover the entailment label with high accuracy (91.58%, 94.79%, 87.17%, and 99.83%). Mistral provides the sharpest dissociation: its probe reaches 99.83%, while direct behavior is 50.67% and forced True/False choice is 51.50%. At L16, a 10,000-run pair-preserving permutation test on fixed out-of-fold predictions gives $p=0.0001$, and accuracy remains above 98% from PCA-32 through PCA-256. Yet transfer is incomplete: canonical-to-reordered transfer is 36.58%, while reordered-to-reordered recoverability is 88.08%. Text-only TF-IDF does not reliably recover the label, and steering effects are small and model-dependent. These results support treating recoverability, behavioral expression, transferability, and causal controllability as distinct empirical properties.
Chat is not available.
Successful Page Load