Conflict-Aware Logit Adapters for Utility-Preserving Anti-Distillation
Abstract
While reasoning traces improve model capability, they also create a pathway for extracting supervision signals for downstream distillation. Initial efforts to mitigate this rely on decoding-time interventions, but they incur high inference overhead and often harm the teacher model's natural fluency. We propose Conflict-Aware Logit Adapters (CALA), a lightweight training-time defense. CALA attaches a small residual adapter to a frozen teacher model and trains only this adapter. A proxy student identifies positions where imitation is most sensitive, and the adapter perturbs the output distribution there within high-probability regions. A conflict-aware optimization balances competing anti-distillation and utility objectives, with regularization preserving the teacher’s performance and fluency. At inference, CALA adds negligible overhead with no auxiliary models required. Experiments on reasoning benchmarks show that CALA substantially reduces successful student imitation while maintaining near-original teacher performance and generation quality, offering a practical and efficient alternative to decoding-time approaches.