A Surrogate Perspective on Convergence of Fixed-Target DQN
Zichu Liu ⋅ Nneka M Okolo ⋅ Ryan D'Orazio ⋅ Danilo Vucetic ⋅ Ioannis Mitliagkas ⋅ Gauthier Gidel
Abstract
Deep Q-Networks (DQNs) approximate value iteration by freezing a target network, forming Bellman-optimality targets, and running an inner regression loop that minimizes a smooth \emph{surrogate loss} (e.g., MSE/Huber loss) on data from a replay buffer. Yet, control is governed by the Bellman optimality operator's $\gamma$-contraction in $\ell_\infty$, whereas common objectives (MSE/Huber loss) measure progress in an averaged $\ell_2/\ell_1$ geometry; thus, improving the surrogate need not reduce the worst-case Bellman error that drives performance. To bridge this gap we connect the progress in the inner loop for each surrogate to the Bellman residual under the $\ell_\infty$ norm. Our analysis provides explicit thresholds denoted by $\alpha$ under which sufficient progress guarantees a $\ell_\infty$ contraction up to a standard approximation floor. This yields new convergence guarantees for DQN with fixed targets trained using MSE or Huber-type losses, and motivates a novel soft-$\ell_\infty$ surrogate that smoothly approximates the sup-norm and better matches the $\ell_\infty$ geometry resulting in the least stringent inner loop accuracy requirements. Across Atari benchmarks, soft-$\ell_\infty$ consistently reduces the worst-case Bellman residual and matches or improves returns relative to standard objectives, providing a practical, geometry-aligned recipe for stable DQN training.
Chat is not available.
Successful Page Load