AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
Abstract
Self-distillation enables language models to learn in an on-policy fashion from their own trajectories by using the same model as both the teacher and student, with the teacher being conditioned on privileged information that is unavailable to the student. This information can take multiple forms or views, including solutions, demonstrations, or feedback; for example, for a math problem, one view might be a ground truth reasoning trace. Supervising the student via the privileged teacher enables fine-grained token-level feedback and removes misalignment caused by distillation from external policies with different semantic biases. However, self-distillation also creates a fundamental asymmetry: the teacher may encode view-specific information tied to the privileged information provided during training, which the student cannot access at inference time. Moreover, the best form of privileged information is often task-dependent, making it difficult to choose a single teacher view. In this work, we address both these challenges jointly by introducing AVSD (Adaptive-View Self-Distillation), a novel method of self-distillation with multiple privileged-information views, which reconstructs token-level supervision by separating stable cross-view consensus from view-specific residual support. The consensus signal provides a robust update direction, while the residual signal is selectively used to adjust the update magnitude. Together, these components form a gated aggregation mechanism that balances a conservative target (geometric mean) with a more permissive target (arithmetic mean) for aggregating teacher signals at the token level. Experiments on math competition benchmarks (AIME24, AIME25, and HMMT25) show that AVSD consistently outperforms single-view self-distillation baselines and GRPO, achieving 3.1\% Avg@8 gain over the strongest baseline on Qwen3-8B on average. Moreover, on code-generation benchmarks (Codeforces, LiveCodeBench) using Qwen3-8B, AVSD outperforms single-view self-distillation baselines by 2.9\% on average. More broadly, our analysis shows that AVSD provides better token-level learning signals than any single-view method by reconstructing supervision from multiple privileged teachers.