Predict, Observe, Retrain: Can a Language Model Learn to Pick A/B Test Winners from Its Own Feedback Loop?
Abstract
Companies spend a lot of money running A/B tests on headlines, ads, and email subject lines. A model that could predict the winner before the test runs would save much of that money. Recent work suggests large language models (LLMs) can simulate human responses, but most evaluations are one-shot: the model predicts, and that is the end of it. Real deployment would be a loop. The model predicts, the real test runs anyway, the outcome comes back, and the model retrains before predicting the next batch. This paper runs that loop on real data. We take 2,599 headline A/B tests from the Upworthy Research Archive (43 million impressions of real reader behavior), sort them into ten time-ordered rounds, and require every model to predict each round's winners using only outcomes revealed in earlier rounds. We compare a frozen zero-shot LLM, a few-shot prompted LLM, a gradient-boosted trees baseline on text features, and a small LLM (Qwen2.5-3B) that is LoRA fine-tuned from scratch each round on all outcomes revealed so far. The full loop was run twice with independent random seeds. The fine-tuning loop wins clearly. It picks the winner of decisive tests 52.8% of the time on average across the two runs (53.7% and 52.0%) against a 25.6% random baseline, beating the best non-LLM baseline by 15 points in both runs and in every round. The gain arrives almost entirely with the first round of feedback and holds steady, rather than climbing round over round. Zero-shot and few-shot prompting barely beat the trees baseline, which matches earlier reports that prompting alone captures little about real preferences. Two checks back the result up. A memorization probe finds the base model cannot reproduce a single archive headline verbatim (0 of 200). The fine-tuned model also transfers: trained only on Upworthy tests, it ranks 1,803 news headlines from the Microsoft MIND dataset with Spearman correlation 0.37 against real click-through rates, where the base model manages 0.04. The loop works and the gains arrive early. What the model learns looks like general clickability, not one site's quirks. At 53% accuracy on decisive pairs, though, it is a tool for pruning candidate pools. It does not replace testing.