Modeling quantum neural network gradient with reinforcement learning
Nhan Luu ⋅ Trung D Luu ⋅ Ngoc Nam Pham ⋅ Thang C Truong
Abstract
Variational quantum algorithms offer a promising route to practical quantum advantage on near-term hardware, yet training quantum neural networks (QNNs) remains hampered by two compounding difficulties: the exponential vanishing of gradient variance known as the barren plateau, and the $\mathcal{O}(L \cdot 2^n)$ time–memory cost of differentiating through an $n$-qubit, $L$-layer circuit. We propose RLQ-Grad, a reinforcement-learning-based optimizer in which a classical policy $\pi_\phi$ (a spectrally-normalized PPO agent) learns to propose parameter updates directly, conditioned on the QNN's current parameters, loss, accuracy, and previous update. Because the surrogate gradient is emitted by a classical network rather than obtained by differentiating through the unitary $U(\theta)$, its variance is not constrained by the barren plateau concentration bound, and its per-step cost scales independent of the Hilbert-space dimension. We prove these two properties formally and verify them empirically on a hardware-efficient ansatz across four supervised benchmarks. RLQ-Grad preserves a near-flat gradient-variance curve where backpropagation, parameter-shift, and adjoint differentiation decay by 1–2 orders of magnitude, attains $128\times$–$1841\times$ wall-clock speedups and constant $\sim0.1$ MB memory at $n=12$ qubits, and improves top-1 validation accuracy by up to $+10\%$ over the strongest gradient-based baseline at every circuit scale tested.
Chat is not available.
Successful Page Load