Bias-Variance Optimized Preference Optimization for Large Reasoning Models
Abstract
Large reasoning models (LRMs) generate reasoning traces before producing final answers, yielding strong performance on multi-step reasoning tasks. However, preference alignment for LRMs remains underexplored. The ideal answer-level preference objective marginalizes over reasoning traces, but this trace-marginalized objective is computationally intractable. As a result, practical methods rely on single-trace surrogates, which can induce high gradient variance due to stochastic trace sampling. We cast this challenge as a gradient-estimator design problem and propose Bias-Variance Optimized Preference Optimization (BVPO), which combines a standard trace-based gradient estimator with an empty-trace gradient estimator from a separately sampled empty-trace branch. Under our fixed-prompt, fixed-answer conditional view, the empty-trace branch is deterministic with respect to the trace randomness studied in our analysis. We show that the composite estimator contracts the conditional reasoning-trace-noise component, admits a conditional MSE-optimal mixture, and appears in the estimator-dependent term of a standard biased-SGD bound. Empirically, BVPO improves alignment over baselines on alignment benchmarks, including AlpacaEval 2 and Arena-Hard. Although trained on general conversational data, BVPO also generally preserves reasoning performance on six math reasoning benchmarks. Empirical diagnostics further show that BVPO yields lower total gradient variance during training.