Benchmarking Exploration Hacking Propensity in Frontier LLMs
Abstract
Reinforcement learning (RL) is central to frontier-model post-training and capability elicitation in safety evaluations. Prior work shows that models can exploit RL’s dependence on exploration by strategically altering their exploration to influence training outcomes. This behavior is known as exploration hacking. In capability evaluations, such behavior could undermine the validity of the evaluation. We introduce Exploration Hacking Propensity Bench (EHP-Bench), a scaffolded evaluation of the conditions under which models strategically underperform during capability elicitation. Across six levels, the benchmark varies the information models receive about the consequences of their performance and the motivational pressure to act on it. We evaluate seven frontier models on WMDP-Cyber across single-turn, sequential, and agentic-discovery settings, with five additionally evaluated on WMDP-Bio and Chem in the single-turn and agentic settings. We find that (1) unnudged underperformance is rare and strongly dependent on model and delivery format. No significant single-turn effect appears in any of the seven models, while Gemini~3.1 Pro shows a strong agentic effect and Claude Opus~4 a smaller behavioral effect in the sequential setting. (2) The Gemini~3.1 agentic effect replicates across Cyber, Bio, and Chem and is accompanied by visible strategic reasoning. (3) Stronger scaffolding broadens underperformance, with preference nudges affecting several models and explicit instructions affecting all seven. Overall, exploration-hacking propensity is uncommon in our benchmark but appears under specific combinations of model, training or deployment consequence, and delivery format. EHP-Bench provides a way to test capability evaluations for these early signs of strategic underperformance.