Choosing LLMs for Strategic Agent Interactions
Arastun Mammadli ⋅ Herbert Woisetschläger ⋅ Jawad Fayaz ⋅ Shiqiang Wang
Abstract
Game-theoretic benchmarks claim that large language model (LLM) strategic decision-making varies across game types, interaction structures, and capability dimensions. However, they only evaluate strategic performance across individual fixed models. In parallel, contemporary LLM routers pick models under quality-cost trade-offs, but with no notion of a strategic structure. They do not account for an opponent that might be unknown, and that can adapt over time and shift the feedback signal. To address this, we extend the game-theoretic evaluations and define strategic selection as the meta game layered on top of the base game. In each round, a strategic selector fields an LLM from a portfolio of models, to navigate a game payoff matrix that is latent and learned online. We find that the evaluation signal is game-specific, and that extracting a reliable signal requires a much higher number of seed runs $N$ compared to traditional benchmarking setups. Evaluating across 5 games and 10 LLMs, we show that a model's strategic competence varies significantly depending on the game setup and the opponent faced. Building on this, we implement a suite of selector methods and find that approaches relying on query embeddings or LLM self-reflection do not consistently capture the underlying strategic signal, while a lightweight multi-armed bandit algorithm performs more reliably across games, though still falling short of the best fixed model. Finally, we show that against adapting opponents, simple non-stationary and adversarial bandit variants fail to outperform a stationary bandit baseline, pointing to a gap in how current methods handle a shifting opponent in strategic settings.
Chat is not available.
Successful Page Load