Self-Improvement with Best-of-N and Retraining: Optimal Scaling for Multiclass SGD
Alireza Mousavi-Hosseini ⋅ Murat Erdogdu ⋅ Adel Javanmard
Abstract
We study Best-of-$N$ (BoN) sampling when a model scores its own answers using their log-probabilities. In a multiclass logistic regression problem, we derive the BoN accuracy early in training, and demonstrate that overoptimization, where BoN can overfit the self-reward, can happen even when the data distribution is noiseless. We further demonstrate how input feature correlations and a power-law prior over labels affect the optimal $N$. We then study retraining the model on BoN generated responses, with the goal of improving test-time computational efficiency by sharpening the model around high-likelihood labels. Under a compute budget, we derive the optimal number of samples in BoN, and the optimal number of training steps, for retraining.
Chat is not available.
Successful Page Load