Quantum Reinforcement Learning with Dynamic-Circuit Qubit Reuse and Grover-Based Trajectory Optimization

  • 2025-09-19 14:11:35
  • Thet Htar Su, Shaswot Shresthamali, Masaaki Kondo
  • 0

Abstract

A fully quantum reinforcement learning framework is developed that integratesa quantum Markov decision process, dynamic circuit-based qubit reuse, andGrover's algorithm for trajectory optimization. The framework encodes states,actions, rewards, and transitions entirely within the quantum domain, enablingparallel exploration of state-action sequences through superposition andeliminating classical subroutines. Dynamic circuit operations, includingmid-circuit measurement and reset, allow reuse of the same physical qubitsacross multiple agent-environment interactions, reducing qubit requirementsfrom 7*T to 7 for T time steps while preserving logical continuity. Quantumarithmetic computes trajectory returns, and Grover's search is applied to thesuperposition of these evaluated trajectories to amplify the probability ofmeasuring those with the highest return, thereby accelerating theidentification of the optimal policy. Simulations demonstrate that thedynamic-circuit-based implementation preserves trajectory fidelity whilereducing qubit usage by 66 percent relative to the static design. Experimentaldeployment on IBM Heron-class quantum hardware confirms that the frameworkoperates within the constraints of current quantum processors and validates thefeasibility of fully quantum multi-step reinforcement learning under noisyintermediate-scale quantum conditions. This framework advances the scalabilityand practical application of quantum reinforcement learning for large-scalesequential decision-making tasks.

 

Quick Read (beta)

loading the full paper ...