Source-linked AI summary

Barren plateaus in quantum neural network training landscapes

Jarrod R. McClean, Sergio Boixo, Vadim N. Smelyanskiy, Ryan Babbush, Hartmut Neven

arXiv:1803.11173v1quant-phcs.LGphysics.chem-ph

TL;DR

The paper asks whether random parameterized circuits are suitable initializations for hybrid quantum-classical algorithms. It analyzes gradient concentration using random-circuit and unitary-design properties, finding that gradients vanish with exponentially high probability for sufficiently random circuits. The authors conclude that structured initialization or other mitigation strategies must be studied for scaling beyond a few qubits.

  • Problem

    Random circuits are attractive initial guesses for hybrid quantum-classical algorithms, but their suitability is limited by Hilbert-space concentration and the cost of estimating small gradients.

  • Method

    The paper analyzes random parameterized quantum circuits using Haar-measure properties and unitary 1- and 2-design assumptions, supported by numerical simulations of layered random circuits.

  • Results

    For a wide class of random quantum circuits, gradient averages and spreads concentrate toward zero, producing barren plateaus as circuits approach 2-design behavior.

  • Takeaways & Limitations

    Randomly initialized circuits of sufficient depth offer relatively little utility for hybrid quantum-classical algorithms, so alternative ansatz strategies must be studied.

  • Takeaways & Limitations

    The analysis assumes independent circuit halves in which U−, U+, or both match the Haar distribution through the second moment.

Abstract

from arXiv · show

Many experimental proposals for noisy intermediate scale quantum devices involve training a parameterized quantum circuit with a classical optimization loop. Such hybrid quantum-classical algorithms are popular for applications in quantum simulation, optimization, and machine learning. Due to its simplicity and hardware efficiency, random circuits are often proposed as initial guesses for exploring the space of quantum states. We show that the exponential dimension of Hilbert space and the gradient estimation complexity make this choice unsuitable for hybrid quantum-classical algorithms run on more than a few qubits. Specifically, we show that for a wide class of reasonable parameterized quantum circuits, the probability that the gradient along any reasonable direction is non-zero to some fixed precision is exponentially small as a function of the number of qubits. We argue that this is related to the 2-design characteristic of random circuits, and that solutions to this problem must be studied.

Gradient concentration in random circuits

The paper studies random parameterized quantum circuits whose layers combine parameterized unitaries with non-parameterized unitaries, under assumptions linking circuit halves to Haar-random second moments. These assumptions make gradient averages vanish and cause gradient variance to concentrate around zero as circuits approach unitary 2-designs.

  • Circuit model: The studied circuit layers use a parameterized unitary Ul(θl) followed by a generic angle-independent unitary Wl.The parameterized component is Ul(θl)=exp(−iθlVl), with Vl Hermitian.
  • Circuit model: Random circuits are modeled so that, for every gradient direction, U−, U+, or both are independent and match the Haar distribution through the second moment.This is the paper’s defining randomness assumption for the random parameterized quantum circuits.
  • Design assumptions: Unitary t-designs reproduce Haar averages for polynomials of degree at most t in U and U∗, while exact Haar invariance generally requires exponential resources.The paper uses this framework to connect circuit depth and gradient statistics.
  • Gradient concentration: The gradient average is zero when the relevant circuit halves are independent and at least one forms a unitary 1-design.The distribution p(U) specifies how circuit unitaries are sampled.
  • Gradient concentration: Levy’s lemma explains why Haar-random n-qubit states produce exponentially decreasing measurement variance as the Hilbert-space dimension grows.The states occupy a hypersphere of dimension D=2^n−1, and the gradient derivative is Lipschitz continuous.
  • Gradient concentration: When either circuit half forms a unitary 2-design, the variance matches the corresponding Haar behavior and gradients are likely to lie on a barren plateau.The variance depends on degree-2 polynomials in U and U∗, so a 1-design fixes the average but not necessarily the variance.

Numerical simulations

Numerical simulations show that gradient statistics in modest-depth random circuits concentrate and decay exponentially with system size, while increasing circuit depth drives the variance toward a qubit-dependent 2-design limit.

  • Numerical simulations: The simulations use 1D random circuits whose layers combine random single-qubit rotations with nearest-neighbor controlled-phase gates.The circuit begins with H gates and an RY(π/4) layer; the number of parameters equals qubits times layers.
  • Numerical simulations: Exponential decay is observed in both the expected gradient and its spread as the number of qubits increases for a two-local Pauli term.Figure 3 presents this behavior on a semi-log plot, even when the layer count is a modest linear function of qubit number.
  • Numerical simulations: As circuit depth increases, the gradient variance converges toward a fixed second moment determined by the number of qubits, indicating transition to 2-design-like statistics.Figure 4 varies layers for even qubit counts from 2 to 24, with the two-qubit curve highest.
  • Numerical simulations: The resulting plateau has a height determined by the number of qubits and supports the predicted vanishing-gradient behavior in modest-sized random circuits.The objective is a single ZZ Pauli operator on the first two qubits, with the gradient evaluated against the first parameter.

Contrast with gradients in classical deep networks

Quantum and classical neural-network gradients differ in both their scaling variable and estimation cost. These differences make small quantum gradients harder to use in optimization.

  • Contrast with gradients in classical deep networks: Quantum gradients become exponentially small in the number of qubits, whereas classical gradients can vanish exponentially in the number of layers.The paper states that quantum gradients are therefore generally exponentially smaller than classical ones.
  • Contrast with gradients in classical deep networks: Quantum output-state normalization contributes to saturation of the gradient at an exponential scale in the number of qubits.
  • Contrast with gradients in classical deep networks: Quantum gradient estimation costs O(1/ϵ^α), compared with O(log(1/ϵ)) scaling for classical batch-gradient estimation.Here ϵ denotes the target gradient scale in the quantum estimation discussion.
  • Contrast with gradients in classical deep networks: When measurements are far below the quantum estimation limit at the gradient scale, gradient-based optimization becomes a random walk with exponentially small probability of leaving the barren plateau.The paper concludes that gradient descent without an additional strategy cannot circumvent this challenge in polynomial time on a quantum device.

Conclusions

Analytical and numerical results show that random quantum circuits concentrate observables near Hilbert-space averages and gradients near zero. The paper therefore points to structured initialization or pre-training as alternatives requiring further study.

  • Conclusions: For a wide class of random quantum circuits, observables concentrate around their Hilbert-space averages and gradients concentrate around zero.The conclusion is supported by both analytical and numerical results.
  • Conclusions: Randomly initialized circuits of sufficient depth have relatively little utility for hybrid quantum-classical algorithms.
  • Conclusions: Structured initial guesses and segment-by-segment pre-training are proposed as approaches to avoid these landscapes, but alternatives must be studied beyond a few qubits.The paper presents these as possibilities rather than established solutions.

Appendix I

Under the stated random-circuit assumptions, the gradient has zero expectation and exponentially decreasing variance with qubit number, with 2-design behavior determining the variance. Numerical plots also show exponential decay and convergence of the second moment with circuit depth.

  • The expected gradient is zero when the circuit halves are independently distributed and at least one matches the Haar distribution through the first moment.
  • Figure 5 shows exponential decay in the sample variance of the gradient for the all-zero-state projector as qubit number increases.
  • Figure 6 shows the second moment converging with circuit layers to a qubit-number-dependent fixed value in 1D circuits.The plotted circuits contain even qubit counts from 2 through 24.
  • The gradient variance decays exponentially with the number of qubits under the paper’s random-circuit assumptions.

Appendix II

The numerical example uses projection onto an all-zero computational state as the objective function and evaluates gradient variance across qubit numbers and circuit layers.

  • The simulation takes the objective function to be the projector H = |00...0⟩⟨00...0| onto the all-zero computational state.
Loading 1803.11173v1…