Source-linked AI summary

Model-Free Quantum Control with Reinforcement Learning

V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, M. H. Devoret

arXiv:2104.14539v2quant-ph

TL;DR

Quantum-control optimization is limited by model bias and by methods that assume unavailable state information or fidelity access. The paper introduces model-free reinforcement learning with parameterized control circuits and measurement-derived rewards, achieving sample-efficient learning across challenging oscillator-control tasks while retaining experimental relevance.

  • Problem

    Quantum control methods based on system simulations can suffer from model bias, while fidelity-based approaches require experimentally costly averaging.

  • Method

    The framework learns parameterized quantum-control circuits through trial-and-error using binary rewards from experimentally feasible measurements, without modeling the system dynamics.

  • Results

    The method learns high-fidelity oscillator-control protocols with 10^6–10^7 runs for Fock and GKP preparation and 10^7–10^8 runs with Wigner reward, comparing favorably with tomographic verification.

  • Takeaways & Limitations

    Model-free reinforcement learning can operate under quantum uncertainty and scarce observability while adapting control policies to experimentally available controls.

  • Takeaways & Limitations

    Learning larger non-classical states becomes increasingly difficult as sample complexity grows with the target state's Wigner negativity.

Abstract

from arXiv · show

Model bias is an inherent limitation of the current dominant approach to optimal quantum control, which relies on a system simulation for optimization of control policies. To overcome this limitation, we propose a circuit-based approach for training a reinforcement learning agent on quantum control tasks in a model-free way. Given a continuously parameterized control circuit, the agent learns its parameters through trial-and-error interaction with the quantum system, using measurement outcomes as the only source of information about the quantum state. Focusing on control of a harmonic oscillator coupled to an ancilla qubit, we show how to reward the learning agent using measurements of experimentally available observables. We train the agent to prepare various non-classical states using both unitary control and control with adaptive measurement-based quantum feedback, and to execute logical gates on encoded qubits. This approach significantly outperforms widely used model-free methods in terms of sample efficiency. Our numerical work is of immediate relevance to superconducting circuits and trapped ions platforms where such training can be implemented in experiment, allowing complete elimination of model bias and the adaptation of quantum control policies to the specific system in which they are deployed.

I. Introduction

The paper develops model-free reinforcement learning for quantum control under stochasticity and minimal observability, avoiding reliance on system models, quantum-state knowledge, or fidelity access. It demonstrates control of oscillator-based quantum states and logical gates, with compatibility with experimental implementation.

  • Simulation-based quantum control is vulnerable to model bias because performance depends on the accuracy of the underlying system model.
  • Deep reinforcement learning offers model-free trial-and-error optimization without access to the dynamics model or its gradients.
  • Quantum control is challenging because large continuous quantum state spaces are only partially observed through stochastic measurements.
  • The proposed framework uses binary rewards from individual experimental episodes and trust-region policy updates, significantly outperforming widely used model-free methods in sample efficiency.
  • The framework learns unitary and adaptive measurement-based control for Fock, GKP, Schrödinger cat, and binomial code states, as well as logical gates on encoded oscillator qubits.
  • Although demonstrated with simulated measurement outcomes, the developed agent is compatible with real-world experiments.

II. Related work

Prior reinforcement-learning proposals addressed several quantum-control tasks, but commonly reformulated them to avoid directly handling quantum observability.

  • Earlier proposals applied reinforcement learning to state preparation, feedback stabilization, quantum gates, error correction, and sensing.
  • These proposals made quantum control more tractable in simulation by avoiding direct treatment of quantum observability.

A. Markov decision process

A Markov decision process models learning as episodic interaction in which observations guide actions, transitions, and reward-driven policy improvement. The paper situates quantum control among continuous, partially observable, stochastic environments.

  • Markov decision process: An agent perceives an environment through observations, selects actions through a policy, and receives state transitions over discrete time-steps within episodes.
  • Markov decision process: The reward signal evaluates the resulting environment state and trains a policy to maximize expected cumulative return without specifying the optimal actions.
  • Markov decision process: Quantum control belongs toward the difficult end of the spectrum because it combines continuous, partially observable, and stochastic interactions.
  • Markov decision process: In the illustrated quantum environment, a classical neural-network agent controls an oscillator and ancilla qubit, receives binary ancilla measurements, and obtains final-state reward from a reward circuit.

B. Quantum control as quantum-observable Markov decision process

The paper formulates quantum control as a quantum-observable Markov decision process in which parameterized circuits interact with an oscillator-qubit system through binary, state-disturbing measurements. The control map is left unmodeled, while experimentally feasible delayed rewards drive trial-and-error policy learning.

  • Quantum control as quantum-observable Markov decision process: The agent applies parameterized control circuits to an oscillator-qubit environment for T discrete steps, receiving observations and producing circuit-parameter action vectors.
  • Quantum control as quantum-observable Markov decision process: Quantum observations carry at most one bit and can cause random discontinuous state jumps, making the environment minimally observable and stochastic.
  • Quantum control as quantum-observable Markov decision process: Unitary control produces a constant observation, whereas measurement-based feedback uses binary σz outcomes generated according to the Born rule.
  • Quantum control as quantum-observable Markov decision process: The simulated environment omits dissipative-bath coupling, although such effects could be incorporated by extending the Kraus maps with uncontrolled quantum jumps.
  • Quantum control as quantum-observable Markov decision process: Unlike simulation-based optimization, the approach does not model the Kraus map; the experimental apparatus implements it while the agent learns action-reward patterns by trial and error.
  • Quantum control as quantum-observable Markov decision process: A reward circuit implements a dichotomic POVM and produces a binary final reward, while intermediate rewards remain zero because the measurement can disrupt the state.
  • Quantum control as quantum-observable Markov decision process: Fidelity oracles used in other approaches require costly experimental averaging, making them impractical for high-dimensional quantum systems.

C. Solving quantum control through policy gradient reinforcement learning

The policy-gradient approach learns stochastic quantum-control policies from binary rewards despite partial observability and sampling noise. It uses neural-network policies, PPO updates, and broad candidate exploration to seek stable convergence without fidelity estimation.

  • The agent learns a policy over observation histories, parameterized by a neural network that outputs Gaussian action distributions.The stochastic policy supports exploration during training and becomes near-deterministic by selecting its mean action after training.
  • Each experimental run evaluates a different policy candidate with a binary reward, prioritizing exploration of more candidates over precise evaluation of one candidate.This allocation is designed for stochastic quantum environments where reward outcomes provide minimal information per episode.
  • Policy-gradient methods estimate performance gradients from episodic binary rewards, and the work uses PPO to discourage destabilizingly large policy updates.PPO was developed to address sudden collapses observed with high-dimensional neural-network policies.
  • The reward-circuit approach avoids expensive fidelity estimation, but stochastic reward outcomes can misorder policies and introduce incorrect gradient contributions.The method therefore faces a navigation problem in continuous policy space despite using only ±1 measurements.
  • The empirical goal is stable convergence to high-fidelity protocols in challenging state-preparation tasks.The following experiments test whether performance improves without collapse or stagnation.

IV. Results

The results study model-free reinforcement learning with modular oscillator-control circuits, including imperfect controls, adaptive feedback, and encoded logical operations. The framework uses parameterized gate modules such as SNAP and displacement gates.

  • The control circuit is built from the universal SNAP and displacement gate set, while remaining compatible with other parameterized circuit choices.The selected gates have been realized in the strong dispersive limit of circuit QED.
  • The modular circuit uses D†(α) SNAP(ϕ) D(α), with circuit depth scaling linearly with state size and many experimentally achievable states requiring approximately five modules.Modularity is presented as an alternative to direct pulse shaping with fewer parameters and potential robustness to small perturbations.
  • The experiments test whether model-free reinforcement learning can learn high-fidelity protocols with realistic training effort while adapting to imperfect controls and measurement feedback.The study also extends to logical gates on oscillator-encoded qubits.

A. Preparation of oscillator Fock states

The Fock-state experiments use experimentally implementable reward circuits and stochastic policy search to prepare oscillator states. The agent reaches high fidelities across states and outperforms alternative model-free optimizers under the shared sampling budget.

  • The Fock reward circuit implements a target-projector-style binary measurement using a photon-number-selective qubit π-pulse conditioned on n oscillator photons.For the Fock task, this provides a feasible and trustworthy reward implementation in circuit QED.
  • The reward circuit uses a first ancilla measurement to remove residual oscillator–qubit entanglement and a second measurement to generate R = −m2.With ideal SNAP, the first outcome is fixed; in realistic settings it serves the disentangling role.
  • F > 0.99 was achieved for every tested Fock state, including F > 0.999 for |1⟩, using a realistic number of experimental runs.The training used 4 · 10^6 experimental runs, while fidelity was reserved for evaluation rather than training.
  • PPO remains stable under stochastic rewards because policy distributions change only modestly, unlike Nelder–Mead and simulated annealing updates that can change policies drastically.The comparison uses the same total sample size, Mtot = 4 · 10^6, across approaches.
  • Target-projector rewards are highly informative for state certification but require an independently calibrated reward-circuit unitary to avoid biasing the learning objective.The Fock reward is identified as a feasible trustworthy instance, whereas more general target projectors may be experimentally difficult.
  • When target-projector rewards are infeasible, probabilistically sampling trustworthy POVM elements provides a more general reward-measurement strategy.The paper applies such schemes to later state-preparation examples.

B. Preparation of stabilizer states

The stabilizer-state experiments extend reward-based control to grid states using probabilistically sampled stabilizer measurements. The agent prepares challenging finite-energy GKP states despite truncation and depth constraints, although partial-information rewards are less efficient.

  • The GKP target encodes a two-dimensional qubit subspace in an oscillator and is relevant to bosonic quantum error correction and other quantum applications.The stabilizer-state construction uses commuting stabilizer generators for the grid state.
  • The reward scheme samples uniformly between x- and p-direction stabilizer projectors, assigning reward scale Rk = 1 for each direction.For these oscillator stabilizers, the expected reward has no simple fidelity relationship, although the required optimization condition is satisfied.
  • Finite-energy stabilizers are measured with an approximate generalized-measurement circuit based on proposals demonstrated in trapped ions and superconducting circuits.The construction adapts experimentally realized phase-estimation and generalized-measurement ideas to the reward circuit.
  • Preparing narrow grid states requires large SNAP truncation, so the experiments use Φ = 30 and T = 9, corresponding to 288 real control parameters.Increasing the action-space dimension can make learning less stable and efficient.
  • The agent successfully prepares grid states under limited SNAP truncation and circuit depth, evaluated by average stabilizer value and illustrated with Wigner functions.Perfect policies would reach stabilizer value +1, but smaller Δ makes this increasingly difficult under the chosen constraints.
  • Probabilistic reward measurements are generally less efficient because individual reward bits carry only partial information about the state.Quantum nondemolition access to multiple commuting stabilizers could improve the reward signal's signal-to-noise ratio.

C. Preparation of arbitrary states

The paper constructs experimentally feasible Wigner-reward estimators from displaced-parity measurements, enabling model-free preparation of arbitrary target states. Increasing phase-space sampling improves reward resolution and attainable fidelity, while hyperparameter choices affect convergence.

  • Wigner reward: The reward scheme uses a tomographically complete, experimentally feasible measurement strategy so arbitrary state preparation is possible in principle.The approach is demonstrated for circuit-QED Wigner tomography and is extended to characteristic-function estimators for trapped ions and multi-qubit systems.
  • Wigner reward: A single binary displaced-parity measurement per policy candidate provides an unbiased Wigner-reward estimator proportional to target-state fidelity.Importance sampling with P(α) ∝ |Wtarget(α)| minimizes variance and yields equal-magnitude rewards.
  • Training results: The agent prepares a cat state in T = 5 steps and a binomial code state in T = 8 steps using Wigner rewards.The target states are |β⟩ + |−β⟩ with β = 2 and 3|3⟩ + |9⟩, respectively.
  • Training results: Sampling 1, 10, or 100 phase-space points increases reward signal-to-noise and reaches higher fidelity, at the cost of increased sample size.Infinite averaging is expected to approach training with fidelity directly available as the reward.
  • Training results: Convergence speed and saturation fidelity vary substantially with reinforcement-learning hyperparameters, and heuristic optimization offers no rigorous convergence guarantee.Multiple random seeds are used to show that suboptimal trapping is unlikely in the presented examples.

D. Learning adaptive quantum feedback with imperfect controls

The paper trains measurement-based feedback directly with imperfect finite-duration SNAP gates, avoiding the degradation caused by policies optimized under idealized controls. The learned agent achieves high-fidelity adaptive preparation even far from the theoretically optimal pulse-duration regime.

  • Control imperfections: The imperfect gate’s short, broad-spectrum pulse drives unintended number-split transitions, leaving the qubit and oscillator entangled after the operation.Long pulses improve selectivity but increase exposure to ancilla relaxation, creating a control trade-off.
  • Adaptive feedback: Post-selection improves a biased policy from 0.9 to 0.97 at χτ = 3.4 but provides no improvement at χτ = 0.4 because it cannot correct incorrect Berry phases.The scheme also has reduced success rate and poor fidelity on unobserved measurement histories.
  • Adaptive feedback: High-fidelity adaptive strategies are learned at χτ = 0.4 despite finite-duration imperfect SNAP gates, where ideal-SNAP policies suffer severe model-bias degradation.The feedback circuit uses measurement outcomes during each episode, and the agent successfully adapts its actions to those observations.
  • Adaptive feedback: At χτ = 0.4, the selected policy reaches average fidelity F = 0.974, with five high-probability branches each exceeding F > 0.9.Post-selecting history hT = 11111 raises fidelity above F > 0.999.
  • Adaptive feedback: The learned policy concentrates on a small set of measurement branches, finishing preparation in three steps on the two most probable branches and idling afterward.Less probable branches use remaining time to recover from undesired measurement outcomes.
  • Model-free learning: Although the demonstration uses a model of finite-duration SNAP, the agent receives only binary measurement outcomes and remains agnostic to the model producing them.This construction allows adaptation to unknown control errors and supports conversion of neural-network policies into decision trees for low-latency inference.

V. Discussion

The discussion examines sample efficiency, scalability, and the trade-offs of model-free reinforcement learning for quantum control. It reports favorable comparisons with model-free baselines while identifying limits from reward variance, action-space dimensionality, sequence length, and on-policy learning.

  • Sample efficiency: 1.06–1.07 experimental runs suffice for high-fidelity Fock- and GKP-state learning, compared with 3·10^6 and 2·10^7 measurements for tomographic verification, respectively.The authors note that exact heuristic sample complexity remains difficult to quantify.
  • Target state complexity: Wigner-reward sample complexity increases with target-state non-classicality, whereas target-projector reward has variance F(1 −F) and a photon-number-independent measurement bound.For Fock states, Wigner negativity grows as √n, implying O(n) scaling for the Wigner-reward bound.
  • Action space: Higher-dimensional action spaces can slow learning, but the agent may disregard irrelevant dimensions for lower Fock states; NM and simulated annealing degrade more strongly on the same problem.The SNAP-and-displacement action space has dimension |A| = Φ + 2, while the experiments use |A| = 17 for all Fock states.
  • Sequence length: Long sequences remain a scalability challenge, although recurrent architectures such as LSTM and newer Transformer attention mechanisms offer supported strategies for long-term dependencies.The discussion frames sequence length as a setting where machine-learning advances may transfer to quantum control.
  • Learning algorithm: PPO was chosen for simplicity and stochastic stability, but its on-policy data discard makes it less sample-efficient than off-policy methods with replay buffers.The authors identify comparing alternative reinforcement-learning algorithms as an open direction.
  • Hybrid initialization: Simulation-based supervised pre-training can provide a better starting policy for real-world retraining and produced significant speedups in preliminary numerical experiments.This retains real-world retraining while using simulation to initialize the policy.
  • Bias–variance trade-off: Model-free learning removes model bias at the cost of sample efficiency, contrasting with sparse physically interpretable models that use few calibrated parameters.The discussion presents this as a bias–variance trade-off rather than a universally superior strategy.

VI. Conclusion

The paper presents end-to-end model-free reinforcement learning as a feasible approach to quantum control under uncertainty and scarce observability. Numerical experiments show stable, sample-efficient learning from stochastic binary measurements across unitary and adaptive feedback tasks.

  • VI. Conclusion: The framework learns stable, high-fidelity control policies directly from stochastic binary measurement outcomes without averaging away the measurement stochasticity.The demonstrated tasks include unitary control and adaptive measurement-based quantum feedback.
  • VI. Conclusion: The model-free agent is intended for real-world experiments where simulation-based methods are not applicable, including the circuit-QED harmonic-oscillator setting studied here.The conclusion frames the approach as eliminating model bias while adapting policies to experimental systems.

Appendix A: Educational example

The educational example shows how a policy can learn a qubit state-preparation action from binary rewards, while the appendix compares this stochastic exploration with Nelder–Mead and simulated annealing. The RL approach uses sampled Gaussian policy candidates and outperforms these alternatives in sample efficiency under the reported comparison.

  • Problem setting: The example prepares |e⟩ from |g⟩ with a single action, a rotation control circuit, and a σz measurement whose outcome supplies the reward.The known optimal rotation is deliberately discovered without revealing the applied unitary to the agent.
  • Training process: Each epoch keeps the policy fixed while collecting B = 30 stochastic episodes, with Gaussian policy parameters updated through actor–critic training.The initial broad policy supports action-space exploration before concentration around more promising actions.
  • Action-space exploration: Unlike Nelder–Mead, which relies on averaged cost estimates, RL assigns each sampled policy candidate a single ±1 reward and shifts the candidate distribution stochastically.This lets the method operate without accurately estimating any candidate’s cost function.
  • Simulated annealing: Simulated annealing is slower than RL even with direct fidelity access, and its performance drops further when fidelity is estimated from 1000 stochastic reward-circuit runs.Under the stochastic estimator, simulated annealing performs worse than both NM and RL.

Appendix C: Variance of the fidelity estimator

The appendix identifies the variance-minimizing phase-space sampling distribution for fidelity estimation and connects it to online reward-based learning. It also describes extensions to other reward schemes and notes practical limits affecting sampling choices.

  • Variance minimization: P(α) ∝|Wtarget(α)| minimizes estimator variance and yields equal-magnitude rewards that stabilize learning.This choice is optimal when one parity measurement is performed per phase-space point.
  • Online estimation: The optimal distribution is independent of the unknown prepared state, making it suitable for online training when only the target state is known.This result assumes Nm = 1 parity measurement per phase-space point.
  • Alternative sampling setting: When both prepared and target Wigner functions are known, the variance-optimal distribution instead satisfies P(α) ∝|W(α)Wtarget(α)|.This alternative applies to fidelity computation rather than the online setting where the prepared state is unknown.
  • Platform extension: The framework extends beyond oscillators to trapped ions, where characteristic-function measurements can provide reward circuits for symmetric states.The characteristic function is tomographically complete, and the proposed circuit uses conditional displacement operations.
  • Multi-qubit extension: For n-qubit state preparation, stabilizer certification protocols can be converted into probabilistic reward schemes, including a stabilizer reward.The proposed multi-qubit framework also includes a characteristic-function reward for arbitrary states.
Loading 2104.14539v2…