Teachers as Modalities: Shared–Private Factorization for Multi-Teacher Distillation
Abstract
Vision foundation models such as DINO and SigLIP learn complementary representations, with strengths in dense visual prediction and language-aligned semantic recognition, respectively. Distilling multiple such teachers into a single compact encoder is appealing, particularly for resource-constrained deployment. We introduce a teacher-as-modality distillation framework that treats heterogeneous teacher representations as complementary views of the same image and explicitly decomposes them into shared and teacher-private components. We first learn a shared--private factorizer that produces one cross-teacher shared representation and two teacher-private representations. We then freeze the factorizer and train a compact ViT-S student to predict these structured targets, while teacher-specific decoders reconstruct the original teacher feature spaces and preserve SigLIP's global vision--language alignment.