SURGE: Sparse Update Routing via Gated Entropy for Efficient Active Distillation
Tanvir Bhathal ⋅ Allen Shen ⋅ Gabriel Hübner ⋅ Dylan Zhou ⋅ Karthik Lakshmanan
Abstract
Mathematical reasoning models generate long trajectories containing both genuine errors and predictable or already-correct tokens. Standard on-policy distillation (OPD) supervises every token, wasting compute, while naive sparse masking destabilizes shared representations. We propose \textbf{SURGE}, combining: (1) \textit{Dual-Gated Sparse Routing}, which uses step-level PRM verification and token entropy to localize uncertain errors, and (2) a \textit{Decoupled KL Objective}, which aligns active failures while anchoring inactive positions. On \textsc{Competition Math}, SURGE reaches \textbf{68.92\%} accuracy, recovering \textbf{34.7\%} of the teacher--student accuracy gap—\textbf{2.4$\times$} OPD's 14.5\% recovery, while using \textbf{74.9\%} fewer distillation FLOPs than OPD. Across five benchmarks, it consistently outperforms SFT, GRPO, PRM, and OPD while reducing FLOPs by \textbf{50.9\%} and verbosity by \textbf{33.0\%} on average.
Chat is not available.
Successful Page Load