Benchmarking LLM Bidding Agents in Electricity Markets: Hour-Specific Pricing and Experience Effects
Abstract
LLM bidding agents are an increasingly plausible presence in electricity markets, yet no prior work has calibrated LLM bidding prices against hourly competitive and joint-profit-maximizing references in electricity markets, or examined whether accumulated market exposure shifts that position. We develop a benchmark com- bining a controlled uniform-price market, exact hourly economic references, a normalized price index φ (0 = competitive; 1 = joint-profit maximizing), and a shared-history base–stress–return protocol with a no-stress counterfactual. Agents receive full price history and profit-maximizing instructions with chain-of-thoughts reasoning. Two findings from a five-firm test case stand out. First, the peak-demand hour (H14) reaches φ14 = 0.907, far above the hourly pure-Nash-equilibrium up- per bound of 0.569; the daily average ¯φ = 0.330 blends 24 hourly outcomes and conceals this peak-hour concentration. Second, high-load exposure leaves a persistent gap of +0.068 in daily φ and +0.028 at H14 after conditions normalize; the low-load effect dissipates. hese outcomes do not establish15 collusion or equi- librium convergence; they show that prices produced by AI bidding agents depend jointly on hour-specific pivotality and accumulated market history, both invisible to standard aggregate monitoring.