Why Ghost Outputs Teach: Towards a Mechanistic Understanding of Subliminal Learning
Abstract
Subliminal Learning (SL) is a recently identified phenomenon in which a student model acquires downstream task capabilities by matching seemingly unrelated auxiliary outputs from a teacher, despite never observing task labels, task-specific outputs, or the original training data. While recent studies have identified where subliminal signals may reside, the optimization mechanism underlying this phenomenon remains poorly understood. In this work, we provide a mechanistic understanding of SL through the lens of learning dynamics. Specifically, we derive a chained cross-task kernel that explicitly links ghost-output supervision to changes in task predictions through shared backbone representations. Our analysis provides a unified explanation for three central empirical observations in SL: (i) under shared initialization, the transfer operator is Positive Semi-Definite (PSD), explaining why ghost-output optimization aligns the student with the teacher's task objective without explicit task supervision; (ii) the ghost-output dimensionality forms an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic high-entropy inputs act as broadband probes that maximize cross-task kernel overlap, explaining why random inputs consistently outperform structured data for subliminal transfer. Experiments on the canonical ghost-output setting validate all three theoretical predictions, providing the first unified mechanistic explanation of how ghost-output supervision gives rise to SL.