Subliminal Traits Persist Through Distillation.
Abstract
Large Language Models trained on data produced by other models can inherit unintended behavioural tendencies, even when such tendencies are not directly observable in the training data. Prior research has demonstrated this phenomenon for a single distillation step. This study examines the transmission of hidden traits across multiple successive distillation steps, where each model is trained solely on its immediate predecessor's outputs. Experiments run on Qwen2.5-7B-Instruct and Gemma-3-4B-it, with a focus on 3 defined traits, indicate that these traits persist above baseline levels for at least four consecutive distillation steps across all tested scenarios. The persistence of these traits suggests that standard data filtering techniques are insufficient to stop their propagation.