DynamAuction: a reinforcement learning environment for repeated auction with dynamic value
Abstract
A buyer who repeatedly participates in auctions for related items often faces a dynamic value problem: what they win today changes what tomorrow's items are worth. Despite the empirical importance of such effects, especially in display advertising, the algorithmic literature on repeated auctions has overwhelmingly focused on the learning problem—treating each opportunity as a contextual bandit—while sidestepping the planning problem induced by the dependence of current value on past wins. We introduce DynamAuction, a synthetic environment for single-buyer sequential bidding. Each won auction contributes reward through a kernel conditioned on a carry-over context that aggregates the bidder's past wins via a commutative monoid. The environment supports first- and second-price auctions, non-homogeneous Poisson arrivals, delayed feedback, saturation, and recency effects. We also introduce a class of Model Predictive Control algorithms built on top of the myopic-optimal valuation—the optimal bid in a second-price auction if no further opportunity were to follow. This class of myopic-optimal MPC consistently outperforms off-the-shelf deep RL baselines (PPO, SAC, TD3). We release DynamAuction to support RL research on repeated auctions.