Shared Truth: Emergent Truth Properties in Large Language Models via Heterogeneous Injection-based Transfer
Isaiah Freeman ⋅ Joed Ngangmeni
Abstract
Representational Engineering (RepE) aims to expose and steer high-level concept directions in language model activations, and recent work shows that hidden states can be transferred across heterogeneous architectures via trained neural adapters. Whether specific concepts survive such transfer remains an open question. We study one such concept, \textit{truth}, by mapping source-model activations through autoencoder adapters into target-model representation space, varying injection strength $\alpha \in [0, 1]$, and reading the result with the target's native truth probe across 21 cross-family trajectories from four model families, with comparison against a closed-form rigid alignment baseline on a representative subset. We find that ordinal truth ranking is preserved through partial injection (median AUROC near the native ceiling through $\alpha \in [0, 0.7]$) but collapses sharply at full injection ($\alpha{=}1$), with parametric separability dropping $\sim$$14\times$ while rank order partially survives. The apparent accuracy overshoot under partial injection decomposes cleanly into two effects: a large recalibration component dominated by native probe miscalibration ($R^2 = 0.75$), and a small but universal genuine separability gain present in all 21 trajectories. These findings characterize cross-family truth structure as partially compatible, ordinally stable, but non-isomorphic: preserved enough under partial injection that learned and rigid alignment methods both surface meaningful truth signal, but with full replacement representing a hard failure mode that practitioners should treat as a deployment boundary.
Chat is not available.
Successful Page Load