Outside the Purview of the Reward: Capability Drift in Multilingual Post-Training.
Abstract
Post-training rewards cover only part of the behavior a deployed model must retain. We study this gap in a mixture-of-experts model serving English and eight Indic languages. Starting from SFT-Instruct, a supervised checkpoint, we train two reinforcement-learning branches. RL-Base improves MATH accuracy by 36.07 points but loses 11.28 points on XQuAD-Indic. XQuAD-English improves, revealing a language-dependent trade-off within the same task family. Adding Indic instruction-following data and language-aware rewards produces RL-Indic, which improves all four regressing Indic benchmarks relative to RL-Base, at a cost in competition mathematics and code. A closer look at MATH reveals a complementary effect: much of the accuracy gain comes from learning to emit the boxed answer expected by the grader. Post-training can therefore strengthen rewarded conventions while weakening other useful behavior. The results motivate evaluating task performance, language compliance, and answer format separately rather than treating a higher aggregate score as evidence of uniform improvement.