Source-linked AI summary

Noise-Induced Barren Plateaus in Variational Quantum Algorithms

Samson Wang, Enrico Fontana, M. Cerezo, Kunal Sharma, Akira Sone, Lukasz Cincio, Patrick J. Coles

arXiv:2007.14384v6quant-phcs.LG

TL;DR

Noise leaves the scalability and trainability of variational quantum algorithms poorly understood. The paper rigorously analyzes local Pauli noise and proves that gradients vanish exponentially with qubit number when ansatz depth grows linearly.

  • Problem

    How noise affects the asymptotic scaling and training process of variational quantum algorithms remains largely unknown.

  • Method

    The paper rigorously bounds gradients for a generic variational ansatz under local Pauli noise, including QAOA and UCC as special cases.

  • Results

    Gradients vanish exponentially in the number of qubits when ansatz depth grows at least linearly with n, producing noise-induced barren plateaus across the cost landscape.

  • Takeaways & Limitations

    Noise-induced barren plateaus impose exponential scaling on variational algorithms and provide quantitative guidance for reducing circuit depth to potentially avoid them.

  • Takeaways & Limitations

    The paper states that layer-wise training, parameter correlation, gradient-free optimization, and higher-order derivatives do not address noise-induced barren plateaus.

Abstract

from arXiv · show

Variational Quantum Algorithms (VQAs) may be a path to quantum advantage on Noisy Intermediate-Scale Quantum (NISQ) computers. A natural question is whether noise on NISQ devices places fundamental limitations on VQA performance. We rigorously prove a serious limitation for noisy VQAs, in that the noise causes the training landscape to have a barren plateau (i.e., vanishing gradient). Specifically, for the local Pauli noise considered, we prove that the gradient vanishes exponentially in the number of qubits $n$ if the depth of the ansatz grows linearly with $n$. These noise-induced barren plateaus (NIBPs) are conceptually different from noise-free barren plateaus, which are linked to random parameter initialization. Our result is formulated for a generic ansatz that includes as special cases the Quantum Alternating Operator Ansatz and the Unitary Coupled Cluster Ansatz, among others. For the former, our numerical heuristics demonstrate the NIBP phenomenon for a realistic hardware noise model.

I. Introduction … B. General analytical results

The paper establishes that local noise can make VQA gradients vanish exponentially with circuit depth and qubit number, creating noise-induced barren plateaus. It develops a general noisy-ansatz framework, proves cost and gradient bounds, and extends the results to correlated noise, degenerate parameters, and measurement noise.

  • I. Introduction: Local noise causes VQA gradients to vanish exponentially with circuit depth, and exponentially in n when depth scales linearly with n.The paper calls this phenomenon a Noise-Induced Barren Plateau (NIBP).
  • A. General framework: The framework models an L-layer parameterized ansatz with local Pauli noise acting before and after each unitary layer.The ansatz encompasses QAOA, Hardware Efficient ansatzes, and UCC, while efficient Pauli decompositions are assumed for relevant Hamiltonians and observables.
  • B. General analytical results: Theorem 1 upper-bounds each noisy-cost partial derivative as a function of L and n, providing the central trainability result.The theorem applies when local Pauli noise with parameter q acts before and after every layer.
  • B. General analytical results: The noisy cost landscape exponentially concentrates around Tr[O]/2^n when the number of layers L scales linearly with the number of qubits.Lemma 1 presents this concentration as a phenomenon accompanying the gradient bound and highlights its relevance to VQE.
  • B. General analytical results: Under efficient Pauli decompositions, the gradient bound F(n) vanishes exponentially in n when L grows linearly in n, independently of the layer or parameter differentiated.The condition holds for all q < 1 and produces an NIBP across the entire cost landscape.
  • I. Introduction: NIBPs are independent of parameter initialization and cost-function locality, flatten the entire landscape, and differ from noise-free barren plateaus.Noise-free barren plateaus concern exponential decay of gradient variance, whereas NIBPs exhibit exponential decay of the gradient itself.
  • B. General analytical results: Correlated or degenerate parameters do not avoid NIBPs, and the same scaling results extend to additional k-local Pauli noise and correlated coherent noise.The degenerate-parameter bound applies at all points in the cost landscape.
  • B. General analytical results: Measurement noise adds a locality-dependent reduction to cost and gradient bounds, with global observables accelerating exponential decay while local observables leave the scaling unaltered.For local observables with w = 1, measurement noise does not alter the scaling; larger locality increases the decay.

C. Application-specific analytical results

The paper derives application-specific conditions under which local Pauli noise produces noise-induced barren plateaus (NIBPs), focusing on QAOA for optimization and UCC for chemistry. It also argues that HVA and certain quantum machine-learning ansatzes may encounter NIBPs under their depth-scaling requirements.

  • QAOA: QAOA is guaranteed to encounter NIBPs when pkP scales linearly in n, where kP is the problem-unitary depth and p is the number of rounds.This condition follows from the QAOA gradient bounds under local Pauli noise.
  • QAOA: Ω(n) depth is inherently required for some graph problems and can arise from hardware compilation, making QAOA NIBPs possible even before accounting for other costs.The Sherrington-Kirkpatrick model has extensive-degree graphs requiring Ω(n) depth circuits; generic hardware mappings may also incur Ω(n) depth or greater.
  • QAOA: QAOA problems with p scaling as poly(n), or even weakly growing p combined with compilation overhead, may encounter NIBPs when pursuing quantum advantage.Some optimization problems are described as requiring p scaling as poly(n).
  • UCC: For UCC with single and double excitations, an Ω(n2Me) depth implementation under 1-D connectivity causes the relevant upper bound to vanish exponentially.The ansatz uses first-order Trotterization and SWAP networks after Jordan-Wigner or Bravyi-Kitaev mapping.
  • UCC: Generalized UCC requires a Ω(n3) depth circuit and therefore exhibits more prominent NIBPs, while sparse UCC implemented in Ω(n) depth still has NIBPs.These depth requirements refer to first-order Trotterized implementations.
  • HVA and quantum machine learning: HVA may encounter NIBPs because its round count scales linearly in n and its compiled unitaries may also grow with n; certain quantum-machine-learning ansatzes may similarly suffer NIBPs with at least linear circuit depth.HVA generalizes QAOA to more than two non-commuting Hamiltonians, while the machine-learning extension uses generalized cost functions.

D. Numerical simulations of the QAOA · E. Implementation of the HVA on superconducting hardware

Realistic hardware-noise simulations show that QAOA performance deteriorates with increasing rounds or problem size, consistent with noise-induced barren plateaus. Hardware implementation of the HVA likewise exhibits exponential decay of gradients and cost differences under noise.

  • D. Numerical simulations of the QAOA: QAOA simulations use gate-set-tomography noise from the IBM Ourense superconducting qubit device to study MaxCut performance.The simulations compare noise-free and noisy training under a realistic hardware noise model.
  • D. Numerical simulations of the QAOA: The study averages approximation ratios over 100 Erdős-Rényi graphs, using 10 optimization runs per graph and 1000 shots per optimization step.For each graph, the run achieving the smallest energy is selected.
  • D. Numerical simulations of the QAOA: For n = 5, noisy-training performance decreases for p > 2, while noise-free-training performance increases with p.The noisy cost concentrates around Tr[H_P]/2^n as p increases, with evidence of exponential decay in cost value.
  • D. Numerical simulations of the QAOA: For p > 4, evaluating noisy-trained parameters without noise yields a decreasing approximation ratio, demonstrating a noise-induced barren plateau effect.The decline indicates that larger QAOA depth makes noisy training increasingly ineffective even when the final cost is evaluated noise-free.
  • D. Numerical simulations of the QAOA: At n ≥ 8, four noisy-trained QAOA rounds fall below the Goemans-Williamson performance guarantee, while circuit depth grows linearly with n.These observations place the simulations in the noise-induced barren plateau regime.
  • D. Numerical simulations of the QAOA: Although n = 5 and p = 4 achieves performance that is NP-hard to achieve classically, scaling to larger problem sizes or depths loses the possible advantage.The results identify noise as a crucial factor in understanding QAOA performance.
  • E. Implementation of the HVA on superconducting hardware: The HVA implementation on IBM’s 27-qubit ibmq_montreal device uses a transverse-field Ising model, local observable O = Z_0Z_1, and layers L = n − 1.Layers increase linearly with qubit number, and entangling gates are restricted to locally connected qubits to reduce SWAP gates.
  • E. Implementation of the HVA on superconducting hardware: Under noise, HVA final-layer partial derivatives and cost differences vanish exponentially until their variances reach the shot-noise floor.In the noise-free case, both quantities decrease only at a subexponential rate.

III. Discussion · IV. Methods · A. Special cases of our ansatz

The discussion emphasizes that local noise fundamentally limits VQA scalability by producing noise-induced barren plateaus, while the methods section shows that the general ansatz includes QAOA, Hardware Efficient, and UCC constructions.

  • III. Discussion: Noise-induced barren plateaus differ from noise-free barren plateaus because gradients vanish with increasing problem size at every landscape point, not merely probabilistically.Consequently, layer-wise training and parameter-correlation strategies that can help with noise-free barren plateaus do not address NIBPs.
  • III. Discussion: Naïvely increasing gradients cannot remove NIBPs’ exponential scaling because it only increases finite-shot derivative variance without improving landscape resolvability.The argument also extends to error-mitigation strategies implementing affine maps to cost values.
  • III. Discussion: Error-mitigation strategies’ ability to mitigate NIBPs remains an open question left for future work.The discussion specifically notes that most error-mitigation techniques consist only of postprocessing noisy results.
  • III. Discussion: The general results apply to a wide range of VQAs, including QAOA for optimization and UCC for chemistry.The paper identifies Corollaries 2 and 3 as treating QAOA and UCC, respectively.
  • III. Discussion: QAOA can exhibit NIBPs in practical optimization problems because a single MaxCut problem unitary may require an Ω(n)-depth circuit, making constant-round circuits at least linear-depth.This conclusion follows from Theorem 1 for combinatorial optimization problems such as MaxCut on 3-regular graphs.
  • A. Special cases of our ansatz: QAOA uses sequential problem and mixer unitaries, UP(γl) = e^−iγlHP and UM(βl) = e^−iβlHM, starting from an initial state such as |+⟩⊗n.The variational parameters determine how long each unitary is applied and are optimized to minimize the cost function.
  • A. Special cases of our ansatz: The Hardware Efficient ansatz reduces gate overhead and circuit depth by using parametrized and native unparametrized gates from a hardware-specific gate alphabet.The illustrated alphabet uses rotations around the y axis and CNOTs.
  • A. Special cases of our ansatz: UCC estimates molecular ground-state energies by applying a parametrized excitation ansatz to a Hartree–Fock reference state, with fermionic operators mapped to spin operators for implementation.Single and double excitations are represented through coupled-cluster amplitudes and excitation operators, with Jordan–Wigner or Bravyi–Kitaev transformations enabling the circuit form.

B. Proof of Theorem 1 · C. Proof of Proposition 1

The proof of Theorem 1 analyzes how alternating unitaries and local noise contract Pauli coefficients and uses this contraction to bound cost-function gradients. The proof of Proposition 1 models measurement noise with local bit flips, equivalently local depolarizing channels, yielding additional locality-dependent damping in gradient and concentration bounds.

  • B. Proof of Theorem 1: Theorem 1’s supporting results use analogous steps, while Corollaries 1, 2, and 3 follow directly from Theorem 1 and Remark 1.The paper provides detailed proofs of Lemma 1 and Remark 1 in supplementary notes.
  • B. Proof of Theorem 1: Theorem 1’s proof tracks operator evolution through Pauli-basis transfer matrices and sandwiched 2-Rényi relative entropy.These tools capture coefficient contraction and movement toward the maximally mixed state under noise.
  • B. Proof of Theorem 1: Local noise multiplies Pauli coefficients according to their X, Y, and Z content, while unitary layers preserve their Euclidean coefficient norm.Together, successive unitary-noise layers shrink the Pauli coefficient vector.
  • B. Proof of Theorem 1: The gradient proof rewrites the output-state derivative and bounds its terms using matrix Hölder, commutator relations, quantum Pinsker’s inequality, and the contraction lemma.The resulting bound is essentially of the form stated in Theorem 1, with the full proof deferred to supplementary information.
  • C. Proof of Proposition 1: Proposition 1 models measurement noise as a tensor product of independent local classical bit-flip channels acting on the POVM elements.The noisy elements are obtained by modifying P0 = |0⟩⟨0| and P1 = |1⟩⟨1|.
  • C. Proof of Proposition 1: Measurement noise is equivalently represented by local depolarizing channels applied directly to the measurement operator.The depolarizing probability is 1 ⩾(1 −qM)/2 ⩾0, as stated in the passage.
  • C. Proof of Proposition 1: Each Pauli coefficient of weight w(i) is damped by qM^w(i), so measurement noise introduces an extra locality-dependent factor into the partial-derivative bound.The operator-norm bound sums these noisy Pauli coefficients before producing the gradient bound.
  • C. Proof of Proposition 1: The same measurement-noise reasoning yields an analogous concentration result for the cost function.This extends the proposition’s measurement-noise analysis beyond the partial derivative bound.

D. Details of numerical implementations … Supplementary Note 1 - Preliminaries

The numerical implementations used tomography-derived IBM Q Ourense noise data and Nelder–Mead optimization, while the supplementary material introduced preliminaries and organized proofs of the main results. The paper also records authors’ roles in conceiving the project and proving its formal results.

  • D. Details of numerical implementations: Gate-set tomography on IBM Q Ourense supplied the one- and two-qubit noise model for numerical simulations.The device was a five-qubit superconducting qubit system.
  • D. Details of numerical implementations: The simulations included native-gate process matrices plus state-preparation and measurement noise described in Ref. [96, Appendix B].
  • D. Details of numerical implementations: The MaxCut optimization used an optimizer based on the Nelder–Mead simplex method.
  • VI. Author contributions: PJC, MC, KS, and LC conceived the project.
  • VI. Author contributions: SW proved Lemma 1 and Theorem 1, while EF and MC proved Proposition 1.
  • VI. Author contributions: KS and SW proved Corollaries 2 and 3.
  • Supplementary Information for Noise-Induced Barren Plateaus in Variational: The Supplementary Information provides proofs for the manuscript’s main results and begins with definitions and lemmas in Supplementary Note 1.Supplementary Note 2 contains a detailed proof of Theorem 1, while later notes continue the supplementary development.
  • Supplementary Note 1 - Preliminaries: Supplementary Note 1 presents preliminaries useful for deriving the results and directs readers to additional quantum-information background references.

1. Definitions

The analysis models a noisy variational ansatz as sequential unitary layers interleaved with local Pauli noise, with costs evaluated as operator expectation values. Quantum states, operators, and layer outputs are represented using Pauli coefficients and vectors.

  • Representation of the quantum state: An n-qubit state is represented by a length-(4n −1) vector of Pauli coefficients ai = ⟨σi n⟩= Tr[ρ σi].The Pauli expansion also applies to Hlm and O, while the identity component is treated separately.
  • Setting for our analysis: The circuit contains L unitary layers, with local noise channels acting before and after each layer on all qubits.The input state is ρ0, and ρl denotes the state after the l-th unitary or noisy-layer application.
  • Setting for our analysis: The noisy cost function is the expectation value of an operator O, and each unitary layer depends on continuous parameters θl with unparameterized gates Wlm.The parameters θ are trained to minimize the cost function.
  • Noise model: Local Pauli channels Nj act independently on each qubit j, with their action on X, Y, and Z determined by qX, qY, and qZ, where −1 < qX, qY, qZ < 1.The noise strength is characterized by a single parameter, as specified in the definitions.

2. Useful lemmas

This section develops lemmas describing how unitary transformations and local Pauli noise affect Pauli coefficients and relative entropy, culminating in bounds for noisy circuits and a commutator identity for gradients.

  • Unitary and noise lemmas: Unitary transformations preserve Schatten norms of Pauli operators and, in particular, the Euclidean norm of their Pauli-coefficient vectors.The invariance follows from unitary invariance of Schatten norms, with the coefficient-norm statement obtained at p = 2.
  • Unitary and noise lemmas: Local Pauli noise contracts every nonidentity Pauli coefficient according to the noise parameters and the Pauli string’s nonidentity support.The relevant support counts satisfy x(i) + y(i) + z(i) ⩽ n and are at least 1 for every nonidentity Pauli string.
  • Relative entropy contraction: Relative entropy contraction is established first for interleaved depolarizing noise and unitaries, then extended to general local Pauli noise with a slightly weaker bound.Depolarizing noise with probability p is a special case of the model with q = (1 −p).
  • Noisy-circuit coefficients: The noisy-circuit lemma bounds the Pauli-coefficient vector after depth l using repeated relative-entropy contraction, unitary invariance, Pinsker’s inequality, and a generic entropy bound.The resulting expression includes the constant c = 1/(2 ln 2) and a depth-dependent factor involving q and l.
  • Gradient lemma: A final commutator lemma gives the relation needed to derive gradient results for unitary gates generated by Hermitian self-inverse operators.It applies to UP(θ) = exp(−iθP/2) for any Hermitian self-inverse P and arbitrary operator A.

Supplementary Note 2 - Proof of Theorem 1

This supplementary note proves Theorem 1 by decomposing the noisy ansatz channel into two parts and bounding the resulting derivative terms. The proof attributes concentration of the partial derivative to state concentrations from depth-L circuits with inserted gates.

  • Theorem statement: Theorem 1 bounds the partial derivative of the noisy cost function for an L-layered ansatz under local Pauli noise acting before and after each layer.The trainable parameter θ_lm corresponds to Hamiltonian H_lm in the unitary U_l(θ_l).
  • Channel decomposition: The overall noisy channel is rewritten as N ◦ W+ ◦ W−, with W− containing l noise layers and W+ containing L−l noise layers.This decomposition separates the channels before and after the layer associated with the differentiated parameter.
  • Derivative bound: The derivative is bounded by applying Hölder’s inequality, using the self-adjointness of the noise channel, and separately upper-bounding two resulting terms.The proof also uses the triangle inequality, Pauli eigenvalues {+1, −1}, and supplementary lemmas.
  • Concentration mechanism: The 1-norm concentration of the partial derivative arises because it can be expressed as a linear combination of state concentrations from depth-L circuits.These circuits include inserted gates Vσj and V†σj between layers, with Pauli coefficients tracked through the output states.
  • Conclusion: Substituting the two term bounds into the preceding derivative inequality yields the theorem’s final upper bound.The proof concludes by inserting equations (86) and (96) into Eq. (83).

a. Stronger bound for low noise levels under more restrictive Pauli noise model · Supplementary Note 3 - Proof of Lemma 1 · Supplementary Note 4 - Proof of Remark 1

The supplementary sections strengthen the noise bound under a restrictive Pauli-channel decomposition, prove concentration for noisy ansätze, and extend the gradient bound to correlated parameters. They also delimit the stronger bound’s narrower applicability while retaining a broader non-trivial result from Theorem 1.

  • a. Stronger bound for low noise levels under more restrictive Pauli noise model: Certain local Pauli channels admit a depolarizing-channel decomposition that yields a stronger bound for low noise and nearly uniform error probabilities.The decomposition uses depolarizing probability p = 4 min(pI, px, py, pz).
  • a. Stronger bound for low noise levels under more restrictive Pauli noise model: The channel is defined by applying I, X, Y, and Z errors with strictly positive probabilities forming a probability vector.Its action is Ppx,py,pz(ρ) = pIρ + pxXρX + pyYρY + pzZρZ.
  • a. Stronger bound for low noise levels under more restrictive Pauli noise model: The stronger construction applies only to a strict subset of Pauli-noise models, whereas Theorem 1 remains non-trivial for the broader class.The modified parameter satisfies q-hat = 1 − 4 min(pI, px, py, pz) ≥ q, with equality for depolarizing noise.
  • Supplementary Note 3 - Proof of Lemma 1: Lemma 1 bounds concentration of the cost function for an L-layer ansatz with local Pauli noise acting before and after every layer.The bound applies to a cost function of the specified measurement-operator form.
  • Supplementary Note 3 - Proof of Lemma 1: The proof represents the pre-measurement evolution as a channel W consisting of L + 1 noise layers interleaved with unitary channels.It then uses the Pauli decomposition, Hölder’s inequality, and Supplementary Lemma 7.
  • Supplementary Note 4 - Proof of Remark 1: Remark 1 extends Theorem 1 to ansätze whose parameters are correlated by being fixed equal to each other.It provides an upper bound on the partial derivative with respect to a degenerate reference parameter at all points in the cost landscape.
  • Supplementary Note 4 - Proof of Remark 1: The correlated-parameter proof counts the g terms in the relevant summation and generalizes directly when those parameters are linear functions of the reference parameter.The reference parameter is the one whose η_lm has the largest 1-norm within the degenerate set.

Supplementary Note 5 - Proof of Remark 2 · Supplementary Note 6 - Proof of Corollaries 2 and 3

Supplementary Note 5 extends the noise model to global Pauli and correlated coherent noise while preserving the main scaling results. Supplementary Note 6 applies the framework to QAOA and UCC ansätze under local Pauli noise.

  • Supplementary Note 5 - Proof of Remark 2: Global unital Pauli noise and correlated coherent noise are proposed as extensions of the noise model.The global Pauli coefficients satisfy −1 ⩽ q_σn ⩽ 1 and q_11⊗n = 1.
  • Supplementary Note 5 - Proof of Remark 2: Under the modified noisy cost function, Lemma 1, Theorem 1, and Corollary 1 remain valid.The modification uses N_i = V_i ◦ P_i ◦ N, with P_i and V_i instances of the extended noise channels.
  • Supplementary Note 5 - Proof of Remark 2: The proof absorbs correlated coherent-noise channels into the parameterized unitaries, reducing the analysis to global Pauli noise.For Pauli coefficients, the argument uses q = max{|q_X|, |q_Y|, |q_Z|} < 1, and the remaining proofs proceed unchanged.
  • Supplementary Note 6 - Proof of Corollaries 2 and 3: QAOA alternates problem and mixer unitaries and contains 2p trainable parameters.The number of Pauli terms in the problem and mixer Hamiltonians is denoted N_P and N_M, respectively.
  • Supplementary Note 6 - Proof of Corollaries 2 and 3: For QAOA, local Pauli noise with parameter q is analyzed when noise acts before and after each native-gate layer implementing the problem and mixer unitaries.Treating each native hardware-gate layer as a unitary layer gives L = (k_P + k_M)p, with the proof invoking Remark 1.
  • Supplementary Note 6 - Proof of Corollaries 2 and 3: For UCC, local Pauli noise with parameter q is analyzed when it acts before and after every U_lm(θ_lm).The resulting statement applies to any coupled-cluster amplitude θ_lm for a molecular Hamiltonian H.
  • Supplementary Note 6 - Proof of Corollaries 2 and 3: The UCC proof uses first-order Trotterization to express the ansatz in the correlated-parameter form required by Remark 1.For the UCC ansatz, the proof identifies g = b_Nlm and ∥η_st∥_1 = 1.

Supplementary Note 7 - Proof of Remark 4

The note generalizes the result to training data encoded in states {ρi} and operators {Oi} of the specified form, then proves it by decomposing the cost into terms Ci and applying Theorem 1.

  • The result generalizes to a training set encoded in states {ρi} and operators {Oi} each of the form (40).
  • Each term in the sum is written as Ci = Tr[OiU(θ)ρiU †(θ)].
  • The proof bounds the resulting expression using Theorem 1.

Supplementary Note 8 - Proof of Proposition 1

Proposition 1 shows that local measurement noise adds a weight-dependent suppression to observable costs and gradients. Consequently, observables with weight w ∈ Ω(n) exhibit exponential decay in n, regardless of circuit depth.

  • Measurement-noise model: Measurement noise is modeled as local bit-flip channels with bit-flip probability (1 −qM)/2, acting on every qubit.The observable is expanded into Pauli strings, and w is the minimum string weight.
  • Measurement-noise model: Basis-independent classical bit-flips correspond to local depolarizing channels applied before measurement.This model also arises when general-basis measurements use noisy one-qubit rotations before computational-basis measurement.
  • Gradient bound: qM^w is the additional factor introduced into the gradient bound relative to Theorem 1.The factor depends on the minimum Pauli-string weight w.
  • Consequence: Observables with w ∈ Ω(n) suffer exponential decay in n of both the cost function and its gradient, independent of circuit depth.The concentration factor is inherited from Lemma 1.
Loading 2007.14384v6…