Strategic Evaluation: Incentivizing AI Capability Coverage with Private Benchmarks
Abstract
Public benchmarks play a significant role in steering LLM development, as market incentives drive model developers to optimize for leaderboard performance. However, fundamental information limitations mean that evaluators are only able to cover a subset of socially relevant capabilities with their evaluation tasks. On the other hand, model developers often have additional private benchmarks that are unavailable to evaluators. The resulting information asymmetry opens the door to strategic task specialization that can lead to suboptimal social welfare. We propose randomized evaluation mechanisms (effectively private benchmarks) as a way to incentivize model developers to train to cover a broader set of socially relevant tasks, including those unknown to the evaluator. We formalize this as a dynamic game with information asymmetry between an evaluator and a model developer. We prove that randomized evaluation aligns incentives to the developer’s prior belief over possible evaluated tasks, and is socially optimal when that belief matches society, but information leakage degrades this over rounds. However, if the evaluator can continually update their known task set, alignment can be asymptotically recovered. We illustrate a variance–leakage–correction tradeoff in semisynthetic experiment with a latent factor structure over MMLU-Pro.