Contamination of Large Language Model Benchmarks in Biology
Arjun Banerjee
Abstract
Benchmarks have long been to be trustworthy mechanism to distinguish different large language model's biological capabilities, but this assumes a level playing field where model's training sets haven't been contaminated with benchmark questions. In this paper, we measure if this assumption still holds. Using 2 contamination detection methods for multiple-choice and open-answer benchmarks across 19 open and closed models, we find that contamination exists in some of the 13 benchmarks we examined. Contamination concentrates on the oldest and most widely copied multiple-choice suites, MMLU-Bio and GPQA-Bio, where the newest models reproduce hidden answers on up to $18\%$ of items, roughly ten times the rate of earlier models. We also detect contamination on recent agentic benchmarks, including ScienceAgentBench and BixBench. Altogether, these results suggest and increasing need to be cautious about public releases of benchmarks and maintaining evaluation integrity when training models.
Chat is not available.
Successful Page Load