Specialist Mediators: Causal Localization of Dense Fine-Tuning
Abstract
Dense fine-tuning can turn a base language model into a domain specialist, but task accuracy alone does not reveal which internal computations mediate the new behavior. We introduce specialist mediators: internal mechanisms through which a dense specialist's lift over its base causally flows. We identify these mediators by patching specialist activations into the base model and asking which patches recover in-domain performance while minimally perturbing off-domain behavior. This separates three questions that are often conflated: where the specialist behavior is localized, how much of it can be recovered by activation patching, and whether a sparse trainable update can implement the same behavior. We formalize these links with exact recovery results, a first-order recovery approximation, and a bound showing when patching recovery can transfer to localized training. In synthetic transformers and structural causal models with planted mediators, patching recovers the true causal mechanisms even when activation-magnitude baselines are misled by non-causal distractors. In Pythia-410M specialists for math (GSM8K), multitask knowledge (MMLU), and biomedical QA (PubMedQA), single-patch localization ranks first on every domain and seed by area under the recovery curve, with small mechanism budgets recovering much of the patchable specialist gap. A localized-LoRA stress test then shows why the separation matters: patching-localized adapters improve negative log-likelihood when the selected attention mechanisms carry the specialist gap, but fail on GSM8K where MLP blocks carry substantial lift. Thus specialist localization is useful both as a route to sparse adaptation and as a diagnostic for when sparse adaptation fails.