When Teachers Imitate Other Models: Identity Transfer to Fine-Tuned Students
Abstract
We find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity claims still follow the producer model. We instruct teacher models (via prompting or few-shot examples) to imitate other models when answering questions. We then fine-tune different student models on their answers and directly probe which identity transfers by asking students for their model name and company. Producer-family identity remains 17.7 percentage points above a human-data control, while identity claims towards the imitated target model change by only +0.28 pp. A DeBERTa classifier that achieves 82.1\% accuracy on held-out teacher answers finds that under cross-imitation both teacher answers (+13.71 pp) and student answers (+10.64 pp) move towards the target. We find that teacher answers reliably predict student writing but this shift does not predict a lift in target-identity claims. Our interpretation is that detectable writing behaviour and verbalized model identity are transmitted through partly different feature channels.