A New Framework for Quantum Reinforcement Learning
Abstract
Quantum reinforcement learning (QRL) has emerged as a promising approach for controlling quantum systems and accelerating learning tasks with quantum resources. However, most existing QRL formulations still retain a classical reinforcement learning interface: states and actions are often encoded in computational bases and rewards are extracted through measurement. This is less natural for fully quantum environments in which the relevant state information may live in an unknown basis and intermediate measurements can disturb the trajectory. This motivates a fully QRL framework in which the agent, environment, rewards, and return accumulation are all represented quantum mechanically. In this paper, we take a step toward this goal and introduce a new QRL framework in which state, action, reward, and reward-accumulation registers evolve without intermediate measurement. In this framework, both the policy and the environment are modeled by conditional quantum channels: the policy updates the action register conditioned on the state, while the environment updates the state and reward registers conditioned on the action, without measuring or overwriting the corresponding control register. We then propose a quantum policy-gradient algorithm for parameterized conditional quantum channels with a trainable control basis. We also derive finite- and infinite-horizon gradient estimators using a new parameter-shift rule. The algorithm samples a single shifted policy block from a geometric distribution, yielding an unbiased gradient estimator under bounded rewards. Finally, we establish convergence of the resulting stochastic gradient-ascent procedure to stationary points under standard smoothness and step-size assumptions.