When Does Fine-Tuning Extract Stored Knowledge? A Relation-Covering Theory for One-Layer Transformers
Abstract
Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. We study this gap in a stylized one-layer self-attention + MLP transformer trained by next-token prediction and subsequently fine-tuned on question-answering data. We first prove that, under suitable regularity conditions, the model reaches near-optimal pre-training loss while learning structured attention patterns. We further show that fine-tuning turns the Q&A prompt format into a trigger for pre-trained relation features, enabling the model to extract facts not revisited during fine-tuning. Our analysis reveals a relation-covering characterization for knowledge extraction: fine-tuning need not revisit every stored fact, but it must cover enough latent relation-template directions through which facts were encoded during pre-training. We prove that extraction improves with pre-training multiplicity and fine-tuning coverage, but becomes harder as the relation-template universe grows. Conversely, insufficient coverage yields a failure regime in which facts can be stored but not extracted, providing a stylized mechanism for hallucination. Our analysis covers both full and low-rank fine-tuning, and experiments on synthetic data and PopQA-based GPT-2/Llama models support the predicted trends.