Skip to main navigation Skip to search Skip to main content

Towards real-time control of a CartPole system on a quantum computer

  • Nguyen Truong Thu Ngo*
  • , Väinö Mehtola*
  • , Jérome Lenssen
  • , Peiyong Wang*
  • , Francesco Cosco
  • , Tien Fu Lu
  • , James Q. Quach*
  • *Corresponding author for this work
  • University of Adelaide
  • Commonwealth Scientific and Industrial Research Organisation (CSIRO)

Research output: Contribution to journalArticleScientificpeer-review

Abstract

The application of quantum reinforcement learning to real-time control systems faces significant challenges regarding hardware latency, noise susceptibility, and learning convergence. This work presents an end-to-end investigation of a minimal hybrid quantum–classical agent applied to the CartPole benchmark, addressing the gap between idealized simulation and execution on a physical superconducting quantum processing unit. We demonstrate that a single-qubit agent acts as an effective learning model, and its performance is functionally comparable to a classical actor-critic network even when the training of the hybrid agent is restricted to use parameter-shift for its quantum circuit component. To connect learning to deployment constraints, we map the inference-time trade-off between control-loop rate and measurement shot budget to provide guidance for an eventual real-time control demonstration. The resulting performance matrices show that both inference control frequency and shot count strongly affect balancing stability: higher inference frequencies consistently improve performance, and increasing the shot budget lowers the minimum inference frequency required to achieve near-maximal balancing. These results highlight the importance of finding an optimal medium between shot count and control frequency and developing circuits that are e.g. initial-state invariant. Lastly, we address the critical bottleneck of control latency on noisy intermediate-scale quantum hardware. By bypassing the standard high-level software stack and programming the Zurich Instruments readout electronics directly via command tables, we achieve more than an order-of-magnitude improvement in execution speed on the VTT Q5 processor. These results quantify some of the current boundaries of quantum-assisted control and provide a start for achieving the tens-of-hertz throughput required for real-time closed-loop control feedback.

Original languageEnglish
Article number045041
JournalMachine Learning: Science and Technology
Volume7
Issue number4
DOIs
Publication statusPublished - 2026
MoE publication typeA1 Journal article-refereed

Funding

J Q Q and P W acknowledges the financial support of the Quantum Technology Future Science Platform—CSIRO.

Keywords

  • NISQ
  • quantum machine learning
  • quantum reinforcement learning
  • real-time control
  • robotics

Fingerprint

Dive into the research topics of 'Towards real-time control of a CartPole system on a quantum computer'. Together they form a unique fingerprint.

Cite this