Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
Zhixuan Liu ⋅ Zhichen Dong ⋅ Yuyu Fan ⋅ Xiangtian Li ⋅ Chao Yang
Abstract
Model distillation can transfer not only intended capabilities but also hidden traits from the teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as $\textit{subliminal learning}$. Prior work establishes this phenomenon, but the transfer mechanism remains unclear, making targeted mitigation difficult. We propose and validate $\textbf{trait-direction drift}$ as the mechanism underlying subliminal learning: $\textbf{(1)}$ on the teacher side, biased teacher samples retain a measurable preference gap toward the target trait even when their semantic content is unrelated to that trait; $\textbf{(2)}$ on the student side, under a low-rank logit-linear approximation, samples with larger preference gaps induce stronger trait-aligned updates during supervised fine-tuning, and these updates accumulate over training. Guided by this mechanism, we propose $\textbf{probe-space corridor}$, a regularizer that constrains drift along a calibrated trait readout during distillation. The method substantially reduces hidden-trait transfer while preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.47% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen-7B preference settings. These results provide both a validated mechanistic explanation of subliminal learning and a targeted recipe for more controllable distillation.
Chat is not available.
Successful Page Load