Knowing Without Saying: How Contextual Evidence Survives but Fails to Surface in Transformers
Abstract
Large language models frequently ignore in-context evidence that contradicts their parametric priors, producing answers that reflect default beliefs rather than the supplied passage, a phenomenon termed knowledge conflict. This failure is commonly attributed to insufficient encoding of the contextual signal. We challenge this explanation. Through systematic, layer-by-layer analysis in a controlled multi-hop question-answering setting, we show that models faithfully encode the context-supported answer in their intermediate representations, yet suppress it in the final layers before generation. We term this late-layer reversal: the representational support for the evidence answer rises through the middle layers, then collapses in the final ~40\% of the network. To explain this phenomenon, we identify a Dual-Track computational structure within the transformer's hidden state. A knowledge direction, recovered by a learned linear probe, encodes the contextually supported answer; a nearly orthogonal output direction, aligned with the unembedding geometry, determines the model's prediction. These two directions are causally severed: interventions along the knowledge direction produce no change in output, whereas targeted suppression of prior-restoring computation along the output direction reliably causes the model to follow the contextual evidence. Our findings hold across nine model-dataset combinations and reveal that knowledge conflict failures arise not because models fail to represent contextual knowledge, but because their generation pathway declines to consult it.