Towards Precise Knowledge Distillation for Large Language Models via Knowledge Probing
Abstract
Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs). Existing distillation paradigms for LLMs can be broadly categorized into black-box and white-box approaches. Black-box methods typically transfer the chains of thought (CoTs) generated by teacher models to student models, yet they fail to convey the intrinsic knowledge embedded within teacher models. White-box methods aim to alleviate this limitation by leveraging internal representations from teacher models. However, they still face two major challenges. First, it is difficult to identify which knowledge is essential for distillation in student models. Second, it is challenging to transfer this key knowledge from teacher models accurately. To tackle these challenges, we propose KPD, a novel Knowledge Probing-based white-box Distillation framework for LLMs, which is designed to precisely distill the knowledge that the student model lacks and can be incorporated into diverse distillation approaches. Specifically, to identify critical knowledge gaps in student models, we introduce an uncertainty-based probing method that extracts tokens prone to student errors as critical knowledge. To locate the corresponding knowledge in the teacher model, we utilize cumulative gradient calculation to probe where such knowledge is stored and then distill the target layers. Extensive experiments across multiple datasets demonstrate that KPD boosts traditional distillation methods by 0.34%-4.56% in Rouge-L scores. Further ablations show that probing both the student and teacher yields more precise distillation. Our code is available at https://anonymous.4open.science/r/KPD-CA0D.