Circuit-Level Knowledge Distillation for Large Language Models
Abstract
Existing knowledge distillation methods supervise only what a student outputs, leaving how it computes those outputs unconstrained, so students may match teacher behavior through entirely different internal mechanisms. We propose Meta Circuit Distillation (MCD), which reframes distillation as the explicit transfer of reasoning circuits, the structured computational pathways uncovered by mechanistic interpretability. MCD represents each model's computation as a transcoder-based attribution graph and aligns teacher and student circuits via an optimal-transport pathway-matching loss that plugs into any standard distillation pipeline. We prove that minimizing this loss bounds the discrepancy in MLP-level computation between teacher and student. Across two model families, six distillation objectives, and three instruction-following benchmarks, MCD delivers consistent improvements that transfer to mathematical reasoning, with the largest gains under aggressive compression.