Source-linked AI summary
Reinforcement Learning in Different Phases of Quantum Control
Marin Bukov, Alexandre G. R. Day, Dries Sels, Phillip Weinberg, Anatoli Polkovnikov, Pankaj Mehta
TL;DR
Preparing quantum states quickly and accurately is difficult in experimentally constrained and computationally complex many-body systems. The paper applies model-free RL, using fidelity-based feedback to construct control protocols, and finds nearly optimal, robust solutions alongside glassy phase structure in the control landscape. The approach remains computationally limited by the action and state-space sizes as system size grows.
Problem
Fast, high-fidelity quantum state preparation is difficult because realistic controls are constrained, decoherence limits time, and large many-body systems are hard to simulate.
Method
A modified Watkins Q-Learning agent constructs driving protocols from fidelity rewards while learning through interaction with simulated quantum dynamics, using exploration and replay.
Results
RL finds stable suboptimal protocols with performance rivaling optimal solutions and identifies nearly optimal variational protocols in glassy control landscapes.
Takeaways & Limitations
Model-free RL can provide useful protocols and insights for nonequilibrium quantum control without requiring accurate system models or explicit local gradients.
Takeaways & Limitations
For a single global drive, the learning part is system-size independent, but the RL algorithm is computationally limited by the action and state-space sizes.
Abstract
from arXiv · showhide
The ability to prepare a physical system in a desired quantum state is central to many areas of physics such as nuclear magnetic resonance, cold atoms, and quantum computing. Yet, preparing states quickly and with high fidelity remains a formidable challenge. In this work we implement cutting-edge Reinforcement Learning (RL) techniques and show that their performance is comparable to optimal control methods in the task of finding short, high-fidelity driving protocol from an initial to a target state in non-integrable many-body quantum systems of interacting qubits. RL methods learn about the underlying physical system solely through a single scalar reward (the fidelity of the resulting state) calculated from numerical simulations of the physical system. We further show that quantum state manipulation, viewed as an optimization problem, exhibits a spin-glass-like phase transition in the space of protocols as a function of the protocol duration. Our RL-aided approach helps identify variational protocols with nearly optimal fidelity, even in the glassy phase, where optimal state manipulation is exponentially hard. This study highlights the potential usefulness of RL for applications in out-of-equilibrium quantum physics.
I. INTRODUCTION
Quantum state preparation is broadly important but difficult under short-time, control, decoherence, and many-body complexity constraints. The paper uses model-free RL to discover high-fidelity protocols and identifies glassy structure in the control landscape.
- Quantum state manipulation supports applications from nuclear magnetic resonance and cold atoms to trapped ions, quantum optics, superconducting qubits, and quantum computing.
- Short, high-fidelity control is difficult because adiabatic assumptions can fail, experimental controls are bounded, and decoherence limits usable evolution time.
- The paper applies modified Watkins Q-Learning to construct time-dependent driving protocols from an initial state to a target state using fidelity as the reward.
- RL is model-free and feedback-based, enabling control discovery when system models are inaccurate or parameters are uncertain, without explicit local gradients.
- The control problem exhibits overconstrained, glassy, and controllable phases as protocol duration increases, while RL finds stable suboptimal protocols that rival optimal solutions.
II. REINFORCEMENT LEARNING
The paper formulates fidelity optimization as a model-free, episodic Q-Learning task in which piecewise-constant controls are selected through interaction with Schrödinger dynamics. Exploration, replay, and repeated initialization help the agent find nearly optimal protocols despite difficult many-body landscapes.
- RL uses a modified Watkins Q-Learning algorithm with linear function approximation and eligibility traces to search for high-fidelity protocols.
- Each episode has NT = T/δt steps, with states containing time and field, actions given by field jumps, and rewards in [0, 1].
- The agent receives final reward Fh(T) = |⟨ψ∗|ψ(T)⟩|2, so intermediate quantum states are not directly included in the optimization objective.
- The environment evolves the quantum state under the time-dependent Hamiltonian generated by the agent’s magnetic-field protocol.
- Q-Learning estimates action values and selects the policy maximizing expected total return, while Bellman solutions remain unavailable in complex non-integrable systems.
- Exploration and replay training expose the agent to broad protocol regions and high-fidelity examples, reducing bad shots and retaining the best encountered protocol.
- The method’s fidelity depends on the initial and target states: small changes cause marginal drops, whereas states from different phases produce a fidelity decrease.
A. Single Qubit Manipulation
The single-qubit control problem has distinct phases as protocol duration changes, with RL identifying physically meaningful protocols and revealing transitions in the optimization landscape.
- RL benchmark: RL finds protocols resembling fast-forward driving, and its model-free approach benchmarks successfully against the analytically solvable single-qubit problem.The initial and target states are ground states at hx=-2 and hx=2, respectively; this choice does not determine RL applicability.
- Control landscape: The single-qubit problem maps protocol optimization onto an infidelity landscape whose local minima are analyzed using stochastic descent and the correlator q(T).The correlator characterizes correlations among local minima and is related to the Edwards–Anderson order parameter.
- Control phases: For T > TQSL ≈2.4, infinitely many unit-fidelity protocols exist in the controllable phase III, making optimization easy.The phase contains exactly degenerate, uncorrelated global minima corresponding to unit-fidelity protocols.
- Control phases: For Tc < T < TQSL, correlated non-degenerate local minima form phase II, where unit fidelity is physically unattainable and optimal-fidelity search becomes difficult.The RL agent discovers a protocol that first brings the state to the equator before directing it toward the target.
- Control phases: For T < Tc ≈0.6, q(T)=0 and a unique solution indicates an overconstrained, effectively convex landscape, although achievable fidelity may remain limited.The transition time tends to zero as the maximum allowed field strength |hx| tends to infinity.
B. Many Coupled Qubits
For coupled qubits, the overconstrained-to-glassy transition persists while the quantum-speed-limit transition moves outside the short-time regime, yet RL finds nearly optimal high-fidelity protocols.
- Model and setup: The coupled-qubit study uses a closed chain of L interacting qubits and the same piecewise-constant control-field setup as the single-qubit problem.Paramagnetic ground states at hx=-2 and hx=2 serve as initial and target states, and many-body fidelity supplies the reward.
- Phase structure: The overconstrained-to-glassy critical point Tc survives in the coupled-qubit model, while TQSL lies outside, or possibly beyond, the short-time range studied.Consequently, the glassy phase extends to long and probably infinite protocol durations.
- Protocol quality: Even without achievable unit fidelity, short protocols include nearly optimal solutions with extremely high many-body fidelity despite exponentially growing Hilbert-space dimension and one control field.The paper cautions that similar fidelities need not imply similar physical properties or resource costs.
- Finite-size behavior: For L ≥6, both q(T) and −L−1 log Fh(T) converge to thermodynamic-limit values without visible finite-size corrections.The authors relate this behavior to Lieb–Robinson-limited information propagation over approximately JT=4 sites at the longest durations considered.
IV. VARIATIONAL THEORY FOR NEARLY-OPTIMAL PROTOCOLS
The RL protocols reveal low-complexity structure in the many-body dynamics, motivating simple variational approximations built from only a few control pulses.
- Variational motivation: The optimal bang-bang protocols keep half-system entanglement entropy small throughout evolution, satisfying an area law.This suggests evolution near the ground state of a local, previously unknown effective Hamiltonian.
- Variational motivation: Their emergent low-entanglement structure motivates constructing simple few-bang variational protocols from the machine-learned solutions.The variational ansatz is intended to capture the nearly optimal protocols identified by RL.
A. Single Qubit
Single-qubit RL protocols exhibit a simple pulse structure that supports a three-pulse variational description, reproducing optimal fidelity and the two critical times up to the quantum speed limit.
- Protocol structure: For T < Tc, the optimal bang-bang protocol has a single jump at T/2, while Tc ≤T ≤TQSL produces multiple bangs around the midpoint.The number of bangs grows with protocol duration in the correlated regime.
- Three-pulse ansatz: A three-pulse ansatz uses symmetric outer pulses separated by an hx=0 interval, with τ^(1) determining the pulse durations.The first pulse brings the state to the equator, and the final pulse directs it toward the target.
- Variational phases: For T ≤Tc, optimization gives τ_best^(1)=T and a zero-duration middle interval, recovering a single-bang protocol.The overconstrained-to-correlated transition appears as a non-analyticity at τ_best^(1)(Tc)=Tc≈0.618.
- Variational phases: For Tc ≤T ≤TQSL, τ^(1) remains fixed while the middle-pulse duration changes, reflecting the equator as the only geodesic for the relevant Bloch-sphere rotation.The simplified ansatz ceases to be valid after the quantum speed limit regime begins.
- Validation: The variational fidelity nearly perfectly agrees with stochastic-descent and optimal-control results and captures both Tc and TQSL.The variational solution coincides with the global infidelity minimum up to the quantum speed limit.
B. Many Coupled Qubits
In coupled many-body systems, a one-parameter variational protocol captures the overconstrained phase but fails in the glassy phase. A two-parameter extension, motivated by RL-discovered protocols, reproduces optimal behavior much deeper into the glassy regime.
- One-parameter theory: The one-dimensional variational ansatz captures the critical point Tc ≈0.4 but breaks down for T > Tc in the glassy phase.Its non-analyticity at Tc is physical, whereas additional kinks can be artefacts of the simplified variational theory.
- Two-parameter theory: The five-pulse extension adds a second independent pulse parameter while preserving protocol symmetry and reducing to the simpler ansatz when τ (2) = 0.The added pulses are reminiscent of spin-echo protocols and appear important for entangling and disentangling the state.
- Two-parameter theory: The two-dimensional variational fidelity seemingly reproduces the optimal fidelity for all protocol durations T ≲3.3 and reduces to the one-dimensional theory for T ≤Tc.Both pulse lengths exhibit a non-analyticity at Tc, with another feature near T′ ≈2.5.
- RL-guided interpretation: The variational theory was derived from RL-discovered protocols and provides a simple effective description of optimal many-body control deep into the glassy phase.The authors connect this strategy to constructing effective theories for out-of-equilibrium physics.
- Two-parameter theory: The two-parameter ansatz is strictly better than the single-parameter protocol for all T > Tc, with its additional pulse controlling and suppressing entanglement entropy during the drive.The effect of the second pulse becomes prominent only at later times, around T ≈1.3.
- Protocol landscape: The density of protocols is compared between the overconstrained phase at T = 0.4 and the glassy phase at T = 2.0 using 1-spin and 2-spin excitations.The comparison uses L = 6 and protocols with NT = 28 bangs.
V. GLASSY BEHAVIOUR
The glassy phase contains a rugged optimization landscape with multiple attractors and low-fidelity local excitations, making optimal control difficult. Nonlocal updates do not remove the underlying glassiness, although they suppress intra-attractor fluctuations.
- Landscape structure: In the glassy phase, 1-flip excitations above the optimal protocol have relatively low fidelities, indicating that local moves do not readily reach the ground state.The protocol density of states is evaluated by enumerating 2^28 bang-bang protocols for L = 6 and NT = 28.
- Multiple attractors: Stochastic Descent at T = 3.2 exhibits three main attractors across 10^3 random initial conditions, each associated with a representative protocol profile.The figure also reports attractor populations and averaged protocol profiles.
- Landscape structure: A 2-flip algorithm would not observe glassiness for T ≲2.5 but would observe it for T ≳2.5, suggesting additional transitions within the glassy phase.The same transition is reflected in the improved two-parameter variational theory.
- Multiple attractors: Protocols within the same attractor can have small mutual overlap, comparable to overlap between different attractors, so moving within an attractor requires highly nonlocal moves.This structure distinguishes the many-body glassy phase from the single-qubit system.
- Algorithmic implications: GRAPE’s nonlocal updates significantly suppress intra-attractor fluctuations compared with Stochastic Descent.The comparison concerns protocol fluctuations within attractors, not elimination of the glassy landscape.
VI. OUTLOOK & DISCUSSION
The paper identifies both the promise and limits of Q-Learning for quantum control, while connecting control-landscape phase transitions to nearly optimal, robust protocols.
- VI. OUTLOOK & DISCUSSION: Q-Learning manipulates single-particle and many-body quantum systems, with performance comparable to leading optimal-control algorithms.The authors identify comparisons among other RL algorithms and Deep RL as future directions.
- VI. OUTLOOK & DISCUSSION: The learning component is independent of system size for a single global drive because it tracks only time, field, and scalar fidelity reward.The computational bottleneck instead arises from action and state spaces, whose current protocol representation limits the algorithm to short protocols.
- VI. OUTLOOK & DISCUSSION: Exact-diagonalization simulations scale exponentially with system size, although Krylov methods can reduce the scaling burden.The number of training episodes also contributes to computational cost because Schrödinger evolution is solved at every time step.
- VI. OUTLOOK & DISCUSSION: The work demonstrates RL suitability for quantum control but does not investigate improvements tailored to quantum-control needs.Suggested directions include alternative state spaces, pre-training, and combining RL with GRAPE or CRAB.
- VI. OUTLOOK & DISCUSSION: RL reveals control phase transitions that also affect state-of-the-art optimal-control methods, indicating universality across optimization approaches.The glassy phase may complicate efficient manipulation and motivate new algorithms informed by the physical interpretation of optimization landscapes.
- VI. OUTLOOK & DISCUSSION: RL-derived variational protocols retain comparable fidelities while being robust to small perturbations and avoiding the exponential complexity of finding the global minimum.The discussion contrasts these protocols with optimal bang-bang protocols and notes RL’s model-free character relative to SGD, GRAPE, and CRAB.
Supplemental Material for:
The supplemental material accompanies the paper “Reinforcement Learning in Different Phases of Quantum Control.”
- The supplemental material is associated with reinforcement learning in quantum control.
- The paper studies different phases of quantum control.
- The supplied title identifies the document as supplemental material for the quantum-control study.
S1. GLASSY BEHAVIOUR OF DIFFERENT MACHINE LEARNING AND OPTIMAL CONTROL ALGORITHMS WITH LOCAL AND NONLOCAL FLIP UPDATES
The supplemental methods describe stochastic descent, CRAB, and GRAPE benchmarks, including their protocol parameterizations, fidelity objectives, gradients, and update rules.
- S1. GLASSY BEHAVIOUR OF DIFFERENT MACHINE LEARNING AND OPTIMAL CONTROL ALGORITHMS WITH LOCAL AND NONLOCAL FLIP UPDATES: Stochastic descent samples bang-bang protocol minima by accepting local field updates that increase fidelity.It restricts fields to hx(t) ∈ {±4} and limits fidelity evaluations to 20×T/δt.
- B. CRAB: CRAB expands the driving protocol in a truncated basis and optimizes its expansion coefficients through a cost function.The implementation uses Fourier harmonics, auxiliary boundary-condition functions, and Nelder-Mead optimization.
- B. CRAB: The CRAB cost combines final fidelity with an L2 protocol penalty to keep optimized controls bounded for comparison.The implementation varies Nc = 10, 20 and selects the best result from 10 random initial configurations.
- C. GRAPE: GRAPE uses gradient optimization for quasi-continuous, piecewise-constant protocols whose magnitudes vary within the allowed manifold.Unlike CRAB, GRAPE is a numeric derivative-based method introduced in NMR spectroscopy.
- C. GRAPE: The GRAPE fidelity gradient is obtained from forward propagation of the initial state and backward propagation of the target state.This procedure requires exact knowledge of the system’s time-evolution operator and therefore an explicit physical model.
- C. GRAPE: GRAPE iteratively updates the control field using the gradient until a desired tolerance is reached.The step size must be adapted because overly large updates can overshoot fidelity maxima or saddles.
- C. GRAPE: The GRAPE step size ϵ_N controls each iteration and should scale appropriately to avoid overshooting landscape extrema.The authors report observing ϵ_N ∝ 1/… in their numerical implementation.
D. Comparison between the RL, SD, CRAB and GRAPE
Across single-qubit and many-body tests, RL reaches competitive or reasonable fidelities while exhibiting the same glassy optimization structure as other methods. The comparison also shows that algorithmic performance depends on protocol duration, control restrictions, and accessible information.
- RL, GS, and GRAPE all find the optimal protocol for the single-qubit system.
- Below the quantum speed limit, CRAB finds good but clearly suboptimal protocols, while increasing its harmonic cutoff does not substantially improve performance.
- In the many-body system with L = 6, all algorithms obtain reasonable fidelities, although the true global minimum uses more bangs than the four-bang variational protocol.
- RL rivals GRAPE at all protocol durations and outperforms CRAB below the quantum speed limit, including the single-qubit glassy phase.
- GRAPE's apparent advantage is qualified because it uses all fidelity gradients and continuous control values hx ∈ [−4, 4], unlike bang-bang RL.
- All algorithms encounter glassiness because the infidelity landscape itself contains many difficult local structures, independent of the optimization algorithm.
- RL also yields better solutions than linear and geodesic protocols in overconstrained and glassy phases, where optimal fidelity remains below unity.
- The optimization landscape changes with protocol duration: controllable phases permit unit-fidelity protocols, whereas glassy and overconstrained phases limit attainable fidelity.