Preserving Exploration for LLM Reasoning via Mean Order-Statistic Alignment
Abstract
Pre-trained large language models exhibit strong exploration ability, as evidenced by high pass@K scores, yet this ability degrades significantly during reinforcement learning (RL) fine-tuning due to mode collapse. A natural remedy is to impose token-wise KL divergence between the fine-tuned and base models; however, we argue that such objective introduces a fundamental mismatch, where token-wise KL operates locally at individual token positions, while exploration is a global behavior that emerges over sequences as a whole. Motivated by this insight, we propose Mean Order-Statistic Alignment (MOSA), which preserves exploration by aligning the fine-tuned model's global behavior to that of the base model. Specifically, MOSA captures global exploration through the mean order-statistic profile, which is obtained by computing the order statistics of each token's posterior over the vocabulary and averaging them across tokens at each rank position. The profile, designed a valid probability distribution, directly admits a principled KL-based objective. For efficiency, we further approximate the profile using only the top-K probabilities and a residual tail bucket, yielding a compact implementation with minimal computational overhead. Experiments on Countdown, 6 mathematical reasoning, and 2 coding benchmarks show that MOSA discovers more diverse plausible solutions, improves both pass@1 and pass@K, and scales effectively to models up to 32B parameters.