Online Evaluation of LLMs via Dyadic Designs
Jinglong Zhao ⋅ Zijie Zhou
Abstract
Online A/B tests are the dominant tool for evaluating new LLM variants in deployment: each user query arrives sequentially, must be routed to one variant before its outcome is observed, and is then scored against the alternative. In practice, these tests almost always use complete randomization. Yet for many queries, simple signals such as predicted difficulty, expected reward, or embedding-based scores are already available, and these are predictive of the outcome. Using them to guide assignment can substantially reduce the variance of the estimated treatment effect. We propose a new dyadic design that uses such covariates to balance treatment and control online across the full covariate distribution, with no hyperparameters to tune. Under stochastic arrivals, our design attains a covariate matching discrepancy of $O(\log^{3} T)$ in one dimension, which is a significant improvement over the $O(T^{1/4})$ rate of existing online stratified designs, and matches the offline minimax rate up to logarithmic factors in higher dimensions, closing a gap left open by prior work. Experiments on a sequential A/B test simulator built from AlpacaEval, in which 805 prompts are scored under two real LLM variants, confirm the theory: our design achieves a covariate matching discrepancy roughly $6\times$ smaller than complete randomization and consistently outperforms the best-tuned online stratified baseline. The method extends to multi-dimensional covariates and to general sequential evaluation pipelines.
Chat is not available.
Successful Page Load