Skip to main navigation Skip to search Skip to main content

Experimental quantum speed-up in reinforcement learning agents

  • V. Saggio (Corresponding author)
  • , B. E. Asenbeck
  • , A. Hamann
  • , T. Strömberg
  • , P. Schiansky
  • , V. Dunjko
  • , N. Friis
  • , N. C. Harris
  • , M. Hochberg
  • , D. Englund
  • , S. Wölk
  • , H. J. Briegel
  • , P. Walther (Corresponding author)

Publications: Contribution to journalArticlePeer Reviewed

Abstract

As the field of artificial intelligence advances, the demand for algorithms that can learn quickly and efficiently increases. An important paradigm within artificial intelligence is reinforcement learning', where decision-making entities called agents interact with environments and learn by updating their behaviour on the basis of the obtained feedback. The crucial question for practical applications is how fast agents learn(2). Although various studies have made use of quantum mechanics to speed up the agent's decision-making process(3,4), a reduction in learning time has not yet been demonstrated. Here we present a reinforcement learning experiment in which the learning process of an agent is sped up by using a quantum communication channel with the environment. We further show that combining this scenario with classical communication enables the evaluation of this improvement and allows optimal control of the learning progress. We implement this learning protocol on a compact and fully tunable integrated nanophotonic processor. The device interfaces with telecommunication-wavelength photons and features a fast active-feedback mechanism, demonstrating the agent's systematic quantum advantage in a setup that could readily be integrated within future large-scale quantum communication networks.
Original languageEnglish
Pages (from-to)229-233
Number of pages5
JournalNature
Volume591
Issue number7849
DOIs
Publication statusPublished - 11 Mar 2021

Funding

We thank L. A. Rozema, I. Alonso Calafell and P. Jenke for help with the detectors. A.H. acknowledges support from the Austrian Science Fund (FWF) through the project P 30937-N27. V.D. acknowledges support from the Dutch Research Council (NWO/OCW), as part of the Quantum Software Consortium programme (project number 024.003.037). N.F. acknowledges support from the Austrian Science Fund (FWF) through the project P 31339-N27. H.J.B. acknowledges support from the Austrian Science Fund (FWF) through SFB BeyondC F7102, the Ministerium fur Wissenschaft, Forschung, und Kunst Baden-Wurttemberg (Az. 33-7533-30-10/41/1) and the Volkswagen Foundation (Az. 97721). P.W. acknowledges support from the research platform TURIS, the European Commission through ErBeStA (no. 800942), HiPhoP (no. 731473), UNIQORN (no. 820474), EPIQUS (no. 899368), and AppQInfo (no. 956071), from the Austrian Science Fund (FWF) through CoQuS (W1210-N25), BeyondC (F 7113) and Research Group (FG 5), and Red Bull GmbH. The MIT portion of the work was supported in part by AFOSR award FA9550-16-1-0391 and NTTResearch.

Austrian Fields of Science 2012

  • 103025 Quantum mechanics
  • 102019 Machine learning

Fingerprint

Dive into the research topics of 'Experimental quantum speed-up in reinforcement learning agents'. Together they form a unique fingerprint.

Cite this