Source-linked AI summary

Experimental quantum speed-up in reinforcement learning agents

Valeria Saggio, Beate E. Asenbeck, Arne Hamann, Teodor Strömberg, Peter Schiansky, Vedran Dunjko, Nicolai Friis, Nicholas C. Harris, Michael Hochberg, Dirk Englund, Sabine Wölk, Hans J. Briegel, Philip Walther

arXiv:2103.06294v1quant-ph

TL;DR

The paper addresses whether quantum communication can reduce reinforcement-learning time, since prior quantum approaches accelerated decisions without demonstrating faster learning. It implements a hybrid agent that alternates quantum amplitude amplification with classical policy updates on an integrated nanophotonic processor, reducing learning time from ⟨T⟩C = 270 to ⟨T⟩Q = 100, a 63% reduction.

  • Problem

    Prior quantum reinforcement-learning demonstrations accelerated action output but did not reduce the average number of agent–environment interactions needed to accomplish a task.

  • Method

    The experiment uses a hybrid agent that exchanges quantum and classical information with the environment, alternating quantum searches with classical policy updates and feedback.

  • Results

    The combined quantum–classical strategy reduces learning time from ⟨T⟩C = 270 to ⟨T⟩Q = 100 for QL = 0.37, corresponding to a 63% reduction.

  • Takeaways & Limitations

    Quantum and classical communication together enable measurable speed-up and control of learning progress, outperforming purely classical communication in the demonstrated setting.

  • Takeaways & Limitations

    Near-term quantum devices have limited coherence times, restricting the number of quantum gates that can be performed before quantum features are lost.

Abstract

from arXiv · show

Increasing demand for algorithms that can learn quickly and efficiently has led to a surge of development within the field of artificial intelligence (AI). An important paradigm within AI is reinforcement learning (RL), where agents interact with environments by exchanging signals via a communication channel. Agents can learn by updating their behaviour based on obtained feedback. The crucial question for practical applications is how fast agents can learn to respond correctly. An essential figure of merit is therefore the learning time. While various works have made use of quantum mechanics to speed up the agent's decision-making process, a reduction in learning time has not been demonstrated yet. Here we present a RL experiment where the learning of an agent is boosted by utilizing a quantum communication channel with the environment. We further show that the combination with classical communication enables the evaluation of such an improvement, and additionally allows for optimal control of the learning progress. This novel scenario is therefore demonstrated by considering hybrid agents, that alternate between rounds of quantum and classical communication. We implement this learning protocol on a compact and fully tunable integrated nanophotonic processor. The device interfaces with telecom-wavelength photons and features a fast active feedback mechanism, allowing us to demonstrate the agent's systematic quantum advantage in a setup that could be readily integrated within future large-scale quantum communication networks.

INTRODUCTION

Reinforcement learning agents improve through feedback, but prior quantum enhancements accelerated decision-making without demonstrating reduced learning time. This work introduces quantum communication between agent and environment and experimentally demonstrates a hybrid approach to speed learning.

  • Quantum mechanics has enhanced reinforcement-learning decision-making, but previous applications retained classical agent–environment communication.
  • The proposed hybrid agent combines quantum amplitude amplification with classical policy updates through a feedback loop.
  • The protocol enables the first quantification and achievement of quantum speed-up in learning time relative to classical communication alone.
  • The experiment uses a fully programmable telecom-wavelength nanophotonic processor with active feedback and potential compatibility with future quantum communication networks.

QUANTUM ENHANCEMENT IN REINFORCEMENT LEARNING

Quantum reinforcement learning reduces learning time by allowing the agent and environment to exchange arbitrary superposition states. A hybrid agent alternates quantum searches with classical tests and policy updates, combining quantum amplification with feedback-based learning.

  • Classical communication restricts exchanges to a fixed preferred basis, whereas quantum communication permits arbitrary superpositions of actions, percepts, and rewards.
  • Alternating quantum and classical epochs creates a feedback loop that makes learning-time speed-up measurable.
  • The framework is designed for near-term devices because only a few or even one coherent query can outperform a classical agent.
  • Classical epochs identify rewards and percepts and update the policy, while quantum epochs search for rewarded action sequences.
  • The quantum-enhanced agent uses the environment as an oracle and amplitude amplification to increase the probability of rewarded action sequences.
  • A quadratic learning-time improvement, ⟨T⟩Q ≤ αJ⟨T⟩C, is achievable when coherent interactions scale with problem size; one coherent query still gives a constant-factor improvement.

EXPERIMENTAL IMPLEMENTATION

The experiment implements hybrid quantum–classical reinforcement learning on a programmable integrated photonic processor. Single-photon circuits realize classical and quantum epochs, with measurements and feedback controlling the learning strategy.

  • The processor comprises 26 waveguides and 88 Mach–Zehnder interferometers in a 4.9 x 2.4 mm integrated platform.
  • Each Mach–Zehnder interferometer uses two thermo-optic phase shifters to act as a fully tunable beam splitter for coherent gate sequences.
  • Photon pairs at telecom wavelength provide a heralded computation photon and a reference photon detected by D0.
  • The experiment encodes action sequences and rewards in two qubits, forming a four-level system represented by processor waveguide paths.
  • The classical strategy measures action outcomes and updates the policy only when detector D2 records a reward.
  • The quantum strategy applies oracle and reflection operations, yielding rewarded-sequence detection probability sin2(3ξ).
  • Agents switch from quantum to classical strategy at ε = 0.396, avoiding amplitude-amplification overshooting.

RESULTS

The experiment compares quantum, classical, and combined learning strategies using average reward and learning time. Switching from quantum to classical operation at the optimal threshold prevents reward decline and substantially reduces learning time.

  • RESULTS: The experiment compares average reward for quantum, classical, and combined quantum-classical learning strategies.Theoretical curves use n = 10,000 agents, while experimental points use n = 165 agents.
  • RESULTS: At ε = 0.396, the average reward begins decreasing for the quantum strategy.This threshold marks the point where the classical strategy becomes more advantageous while quantum iterations continue.
  • RESULTS: The combined strategy switches from quantum to classical operation at ε = 0.396 and always outperforms the purely classical scenario.The comparison is shown in Fig. 4(d).
  • RESULTS: 63% reduction: learning time decreases from ⟨T⟩C = 270 classically to ⟨T⟩Q = 100 with the combined quantum-classical strategy.The winning probability is defined as QL = 0.37, and the reduction agrees well with theoretical values.
  • RESULTS: The combined configuration prevents average-reward decline, while fixed-point alternatives avoid overshooting but provide less favorable speed-up on the limited processor.The reward-control benefit is relevant when Grover search’s intrinsic overshooting drawback is present.

CONCLUSIONS

The paper demonstrates a reinforcement-learning protocol that alternates quantum and classical communication. It reports faster learning and control of the learning process, with photonic technology providing a compact implementation platform.

  • CONCLUSIONS: The protocol alternates information transfer through quantum and classical channels to evaluate performance and control learning progress.Purely classical agents are outperformed in the demonstrated setting.
  • CONCLUSIONS: Integrated photonic circuits provide compactness, tunability, and low-loss communication for reinforcement-learning algorithms.These properties support the platform’s suitability for the demonstrated protocol.

I. Quantum-enhanced learning agents

The paper constructs a hybrid reinforcement-learning agent by combining a classical policy with quantum amplitude amplification. The framework applies to reward-based classical updates and can find rewarded action sequences faster than a matched classical agent.

  • I. Quantum-enhanced learning agents: The hybrid model combines classical reinforcement learning with quantum amplitude amplification through a feedback loop.This approach is designed to determine improvements in sample complexity and learning time.
  • I. Quantum-enhanced learning agents: The agent operates in deterministic, strictly epochal environments where action-percept sequences end with a binary reward.Each epoch begins with the same percept and contains L exchanged action-percept pairs.
  • I. Quantum-enhanced learning agents: The agent’s policy assigns conditional action probabilities and represents behavior within an epoch as action sequences with corresponding probabilities.In deterministic settings, percept histories determine the relevant action probabilities.
  • I. Quantum-enhanced learning agents: Projective simulation supplies the implemented policy, assigning each action sequence a weight factor that is updated after observed rewards.The experiment uses λ = 2, while the broader update method is not limited to projective simulation.
  • I. Quantum-enhanced learning agents: The update can enhance any classical learning scenario when an action-sequence distribution exists and updates depend only on observed rewards.Action sequences are encoded into orthogonal quantum states, and a unitary environment oracle enables quantum search.
  • I. Quantum-enhanced learning agents: The quantum-enhanced agent finds rewarded action sequences faster than a corresponding classical agent with the same initial policy and update rules.The comparison differs by access to the unitary environment oracle.
  • I. Quantum-enhanced learning agents: The optimal Grover-iteration count scales as k ∼ 1/√q when the winning probability q is known or well approximated.Unknown q can be handled by adapting methods for quantum search with unknown reward probability.

I.1. Description of the agent

The hybrid agent alternates quantum amplitude amplification with classical policy updates, choosing between quantum iteration and direct classical sampling according to the estimated success probability. Its expected reward is advantageous below a threshold, after which classical operation should take over.

  • I.1. Description of the agent: The hybrid agent repeatedly performs quantum amplitude amplification followed by a classical epoch and policy update.The cycle prepares an action distribution, applies Grover iterations, tests an action sequence, records reward, and updates the policy.
  • I.1. Description of the agent: The agent performs direct classical sampling when the success probability q falls below the limit Q for which quantum iteration is advantageous.This gives the hybrid agent control over whether the next interaction is quantum-assisted or purely classical.
  • I.1. Description of the agent: For k = 1, the expected rewards are ηC = 2 sin^2(ξ) classically and ηQ = sin^2(3ξ) quantumly.The quantum strategy receives a reward after every second epoch because one epoch is used for the Grover iteration and one for reward determination.
  • I.1. Description of the agent: For q < Q, ηQ exceeds ηC, but at Q = 0.396 the classical strategy begins outperforming continued quantum iteration.The threshold is defined by ηC = ηQ.
  • I.1. Description of the agent: Exact determination of q = ε is not always possible, so Q should be chosen below 0.396 when q is estimated only within a range.Unknown-probability and fixed-point search methods are proposed to determine whether amplitude amplification should be performed.

I.2. Learning time

Learning time is defined by the epochs needed to reach a target winning probability, enabling comparison between classical and quantum-enhanced agents. Quantum search reduces the time to find rewarded action sequences and can yield a quasi-quadratic learning speed-up.

  • Learning-time definition: Learning time T is the average number of epochs needed to reach a predefined winning probability QL; the experiment sets QL = 0.37.The threshold is chosen at or below an intermediate probability Q so that the quantum improvement remains measurable.
  • Learning-time mechanism: Classical and quantum-enhanced agents follow the same policy after finding the same rewarded action sequences, but the quantum-enhanced agent finds those sequences faster.The policy depends on the time-ordered list of observed rewarded sequences, not on whether they were found by classical sampling or quantum amplitude amplification.
  • Classical and quantum times: For a classical agent, tj is the number of epochs needed to find the next rewarded sequence after j − 1 rewards, with average time determined by the successive success probabilities qj.The classical learning time typically scales as ⟨T⟩C ∼ AK for episode length K and A available actions per step.
  • Classical and quantum times: The quantum-enhanced time is quadratically reduced for each search step, with α = π/4 in the reported setting.The parameter α depends on the oracle-query creation cost and on whether qj is known.
  • Learning-time advantage: A quasi-quadratic speed-up in learning time is equivalent to the quantum bound when arbitrary numbers of Grover iterations can be performed.The total bound depends on J, the number of rewarded sequences required for learning, and the classical average learning time.
  • Learning-time advantage: Averaging over possible rewarded-sequence lists with different lengths J still leads to a quadratic speed-up in learning in more general settings.The number of sequences needed to learn can vary when policies or success probabilities depend on the discovered sequences.

I.3. Limited coherence times

Limited coherence restricts how many Grover iterations a near-term quantum agent can perform. The resulting learning-time analysis therefore bounds performance by the available iteration budget and target success probability.

  • Coherence constraint: Near-term quantum devices permit coherent evolution only for a limited time, imposing a maximal number n of Grover iterations.For q = sin2 ξ with (2n+1)ξ ≤ π/2, n iterations maximize the probability of finding a rewarded action.
  • Bounded-iteration analysis: A quantum-enhanced agent limited to n Grover iterations has an analyzed learning time for targets Q < sin2(π/(4n + 2)).This treatment assumes the agent’s policy depends only on the number of observed rewards.
  • Bounded-iteration analysis: For α0 = 1, n >> 1, and (2n+1)ξJ ≪ π/2, the bounded-agent learning time can be approximated using sin x ≈ x.Here sin2 ξj = qj, while α0 sets the epochs needed to create one oracle query.
  • Comparison with the classical strategy: The success probability Qn = sin2(π/(4n + 2)) can be reached in a time whose lower bound is computed from Eq. A.16 with α0 = n = 1, while Eq. A.10 gives the classical strategy.The factor γ depends on the specific setting.

II. Experimental details

The experiment uses telecom-wavelength photon pairs and a programmable silicon photonic processor with thermo-optic control and fast superconducting detectors. These components provide tunable processing and active-feedback capabilities.

  • Photon source: A continuous-wave laser pumps a periodically poled KTiOPO4 crystal in a Sagnac interferometer to generate telecom-wavelength photon pairs by type II SPDC.The crystal is 30 mm long, operates at 25 °C, and is quasi-phase matched for degenerate emission near 1570 nm when pumped near 785 nm.
  • Photon source: The pump wavelength is shifted to 789.75 nm so one photon is produced at 1580 nm for processor coupling and the other at 1579 nm for heralding.The processor is calibrated for 1580 nm.
  • Photonic processor: The processor is a silicon-on-insulator device whose programmable units act as tunable beam splitters controlled by thermo-optical phase shifters.Phase-setting precision exceeds 250 µrad and the phase-shifter bandwidth is around 130 kHz.
  • Detection: Superconducting nanowire detectors provide efficiencies up to 90% at telecom wavelengths, dark counts near 100 c.p.s., timing jitter of hundreds of picoseconds, and reset times below 100 ns.These characteristics support fast photon-detection feedback.
Loading 2103.06294v1…