Learning from Disagreement: Multi-Teacher Distillation for Chinese Spelling Correction
Hongcheng Ding ⋅ Xuanze Zhao ⋅ Ziping Hu ⋅ Hongdan Xiao ⋅ Kang Xu ⋅ Weiyu Zhang
Abstract
Chinese Spelling Correction (CSC) is hard precisely because a single misspelling can originate from any of four loosely coupled dimensions (glyph, phonetic, syntactic, or semantic confusion), yet existing systems treat the four as a flat fusion target. Multimodal pre-training and fixed-weight gating let evidence from each dimension co-exist, but never arbitrate: when a phonetic-plausible candidate contradicts a syntactic-plausible one, the model averages them, and learning effort is spent uniformly on cases where every dimension already agrees. Our diagnosis is that the bottleneck is no longer representation but \emph{decision}: the cases where dimensions disagree are exactly the cases worth specializing for, and dimension-specific teachers expose this disagreement as a usable signal rather than as noise to be smoothed away. We instantiate this view as SEMTD (Self-Evolving Multi-Teacher Distillation), a 4B-parameter student trained from four single-dimension experts (glyph / phonetic / syntactic / semantic) under three coupled objectives: multi-teacher distillation with input-dependent expert selection, self-evolving learning that turns disagreement into a composite reward, and reflective path optimization that replays high-disagreement decisions. Across four benchmarks the student reaches Cor-F1 of $83.8$ on SIGHAN15, $62.3$ on LEMON (7-domain average), $98.2$ on ECSpell (3-domain average), and $73.0$ on CSCD-NS, matching or surpassing 14B LLM-based correctors at $3.5{\times}$ fewer parameters. The takeaway is methodological: in CSC, treating teacher disagreement as the supervision target, rather than a residual to be averaged out, reopens headroom that flat fusion has saturated.
Chat is not available.
Successful Page Load