Source-linked AI summary
Barren Plateaus in Variational Quantum Computing
Martin Larocca, Supanut Thanasilp, Samson Wang, Kunal Sharma, Jacob Biamonte, Patrick J. Coles, Lukasz Cincio, Jarrod R. McClean, Zoë Holmes, M. Cerezo
TL;DR
Barren plateaus make variational quantum optimization exponentially flat and difficult to resolve with finite measurements. This review synthesizes their definitions, causes, architectures, mitigation strategies, and broader connections, while highlighting unresolved questions about useful trainable regions.
Problem
Barren plateaus create exponentially small informative signals, making finite-shot optimization difficult and potentially non-scalable.
Method
The paper reviews recent barren-plateau literature, including causes, architectures, mitigation methods, and connections beyond variational quantum computing.
Results
The review identifies ansatzes, initial states, observables, noise, and circuit structure as factors associated with barren plateaus and surveys methods that can mitigate or avoid them.
Takeaways & Limitations
Barren-plateau research provides guidelines for variational models and insights into quantum information processing, classical simulability, and alternative variational paradigms.
Abstract
from arXiv · showhide
Variational quantum computing offers a flexible computational paradigm with applications in diverse areas. However, a key obstacle to realizing their potential is the Barren Plateau (BP) phenomenon. When a model exhibits a BP, its parameter optimization landscape becomes exponentially flat and featureless as the problem size increases. Importantly, all the moving pieces of an algorithm -- choices of ansatz, initial state, observable, loss function and hardware noise -- can lead to BPs when ill-suited. Due to the significant impact of BPs on trainability, researchers have dedicated considerable effort to develop theoretical and heuristic methods to understand and mitigate their effects. As a result, the study of BPs has become a thriving area of research, influencing and cross-fertilizing other fields such as quantum optimal control, tensor networks, and learning theory. This article provides a comprehensive review of the current understanding of the BP phenomenon.
I. INTRODUCTION
Variational quantum computing combines parametrized quantum models with classical optimization, but scaling can be obstructed by barren plateaus. This review surveys their causes, mitigation strategies, and broader implications.
- Variational quantum algorithms convert problems into optimization tasks trained by classical devices, enabling applications across basic science and machine learning.
- Unlike conventional quantum algorithms, variational schemes are heuristic and can become difficult to train as the number of qubits increases.
- Barren plateaus make loss gradients or differences vanish exponentially with system size, so identifying a loss-minimizing direction can require exponentially many measurement shots.
- Barren plateaus reflect a curse of dimensionality associated with the exponentially large Hilbert space used by variational quantum computing.
- The review synthesizes recent literature into practical guidelines while connecting barren-plateau research to quantum information processing, classical simulability, and alternative variational paradigms.
B. Components of variational quantum computing: data, ansatz, measurements, and classical information processing
Variational quantum computing initializes data, processes it with a parametrized circuit, measures an output quantity, and uses classical optimization to update parameters. Trainability depends on these interacting choices and on finite measurement precision.
- The quantum system starts from an n-qubit state ρ, which may be a fiducial state for VQAs or encode training data for QML.
- A PQC sends the input through a sequence of parametrized unitaries or, with noise or changing qubit numbers, parametrized quantum channels.
- Finite output measurements estimate expectation values or other quantities that a classical optimizer uses to train the circuit parameters.
- In the simplest setting, the loss is the expectation value of one observable O after transforming the input state as ρ(θ) = Uθ(ρ).
- Finite-shot estimation introduces statistical uncertainty, and exponentially small gradients typically require exponentially many shots; training can also face local minima or reachability deficits.
III. TYPES OF BARREN PLATEAUS AND LOSS CONCENTRATION
A barren plateau occurs when losses or gradients become exponentially concentrated as qubit number grows, flattening the optimization landscape. Probabilistic plateaus are average-case phenomena that may retain exponentially narrow useful regions.
- A barren plateau is defined by exponential concentration of the loss or its gradients around their mean as the number of qubits n increases.
- Exponential concentration makes landscapes mostly flat and featureless, so finite-shot optimization can follow statistical fluctuations rather than informative loss changes.
- Probabilistic barren plateaus arise when loss or derivative deviations from their mean become exponentially unlikely for randomly sampled parameters.
- Diagnosing a plateau through the overall loss is more reliable than examining selected partial derivatives, because some derivatives may concentrate while others do not.
- Probabilistic plateaus can contain fertile valleys and narrow gorges, but these useful minima occupy an exponentially small relative volume of parameter space.
B. Deterministic concentration
Deterministic concentration produces a uniformly flat landscape, unlike probabilistic concentration, which can retain narrow well-behaved regions. It can arise from limited state–observable overlap, high-entanglement inputs, or unital noise.
- A deterministic barren plateau bounds the loss deviation from its mean by an exponentially small term for every parameter value.
- Deterministic concentration occurs when the evolved state has exponentially small overlap with the observable across all parameters.
- High entanglement in the input state and unital noise in the circuit are identified as mechanisms that can produce deterministic concentration.
- Unlike probabilistic plateaus, deterministic concentration precludes extrema significantly separated from the loss mean and suppresses all non-exponential landscape features.
- Barren plateaus can be investigated numerically through variance scaling and surrogates, or analytically using assumptions such as unitary designs, local groups, tensor networks, and related methods.
IV. ORIGINS OF BPS
Barren plateaus arise when variational optimization signals concentrate exponentially, linking trainability failures to the exponentially large Hilbert space and circuit-induced exploration of operator space.
- A. A curse of dimensionality: The dynamical Lie algebra generates the reachable unitary group, identifying the circuit transformations available across parameter choices and layers.Its Lie closure contains nested commutators of the circuit generators, and the associated group contains all achievable unitaries.
- A. A curse of dimensionality: Barren plateaus can be understood as a curse of dimensionality caused by exponentially small inner products in an exponentially large operator space.Under general assumptions such as random initialization, the loss inner product becomes exponentially small and concentrated.
- A. A curse of dimensionality: Carefully structured circuits can avoid exponentially small signals by choreographing constructive and destructive interference rather than comparing uncontrolled random vectors.The review connects this possibility to inductive biases and specially designed algorithms.
- A. A curse of dimensionality: Haar averages, t-fold twirls, moment operators, and Weingarten calculus provide tools for evaluating moments of losses over induced unitary distributions.These methods become analytically tractable under assumptions such as group structure and uniform Haar measure.
- A. A curse of dimensionality: The review analyzes how circuit choice, initial states, measurement operators, and hardware noise can each produce exponential concentration.These components determine how the loss landscape explores and compares objects in operator space.
B. Circuit expressiveness
Circuit expressiveness is directly linked to barren plateaus: broader, more unbiased exploration increases loss concentration, while projection structure and finite-depth assumptions determine the precise variance.
- B. Circuit expressiveness: Higher circuit expressiveness makes randomly initialized circuits more prone to exponentially small loss signals and concentrated landscapes.The same relationship has been formalized for both unitary and noisy circuits.
- B. Circuit expressiveness: For deep unitary circuits forming approximate 2-designs, loss variance is determined by the operator-module dimension and the state and observable projection norms.The 2-design condition allows the variance to be analyzed through Haar-like second moments.
- B. Circuit expressiveness: Exponentially large dim(M), exponentially small P_M(ρ), or exponentially small P_M(O) yields exponentially small variance and hence a barren plateau.These factors associate expressiveness, input-state alignment, and measurement alignment with distinct BP mechanisms.
- B. Circuit expressiveness: The 2-design variance formula is exact only under sufficient circuit depth, although related variance bounds remain available for non-2-design and shallow circuits.A BP can still arise in shallow circuits when the adjoint action moves states or measurements through an exponentially large operator subspace.
C. Input states and measurements
Input states and measurement operators can create barren plateaus even when circuit expressiveness is limited, because their operator-module alignment controls the loss variance.
- C. Input states and measurements: Limited circuit expressiveness does not prevent barren plateaus when the measurement operator belongs to an exponentially large module.Small dynamical Lie algebras can nevertheless admit exponentially large operator modules.
- C. Input states and measurements: A polynomial-dimensional measurement module avoids an expressiveness-driven BP, but an exponentially small input-state projection can still cause deterministic concentration.The relevant condition is P_M(ρ) ∈ O(1/2^n).
- C. Input states and measurements: Module purities of the input state and measurement can be interpreted as generalized forms of entanglement, giving operational meaning to their role in BPs.Prior studies separately identify measurement operators and states as determinants of BP presence.
D. Noise
Hardware noise can aggravate or directly induce barren plateaus by enlarging the effective system or driving states toward the maximally mixed state, while some non-unital processes may alleviate the effect.
- D. Noise: Noise channels with the maximally mixed state as fixed point induce deterministic barren plateaus in sufficiently deep circuits.Repeated application drives the state toward the maximally mixed state, suppressing useful optimization signals.
- D. Noise: Non-unital noise and intermediate measurements may alleviate noise-induced BPs, leaving the relationship between noise, measurements, and concentration unresolved.The review presents this relationship as an open question rather than a settled mitigation rule.
- D. Noise: Noise extends the PQC into an environment-augmented space, which can enlarge the effective Hilbert space and aggravate the curse of dimensionality.The noisy loss can be represented using an extended circuit acting on the system and environment.
- D. Noise: Global depolarizing noise drives the loss toward Tr[O]/2^n as circuit depth or depolarizing probability increases.The concentration is explicit in the noisy loss expression for interleaved depolarizing channels.
- D. Noise: For depolarizing probability p ∈ Ω(1/poly(n)), sufficiently large circuit depth can produce a barren plateau.This is a depth-dependent noise-induced BP condition.
A. Hardware efficient ansatz and deep unstructured circuits
Hardware-efficient ansatzes use generic, device-compatible gates rather than problem-specific structure. When sufficiently deep, their universality and expressiveness produce BPs regardless of the initial state or measurement operator.
- Hardware-efficient ansatzes are unstructured circuits of single-qubit rotations interleaved with fixed or parametrized entangling gates.
- Their gates are selected for implementation convenience and device connectivity rather than for a strong problem-specific inductive bias.
- Under general conditions, hardware-efficient ansatzes are universal, with dynamical Lie algebra g = su(2^n).
- Deep hardware-efficient circuits exhibit BPs irrespective of the initial state or measurement operator.
- Problem-inspired ansatzes: Problem-inspired ansatzes encode task structure through gate choices, entangling topologies, or parameter alignment with symmetries, generally reducing expressiveness relative to su(2^n).
- Problem-inspired ansatzes: HVA and QAOA BP behavior depends on the problem: HVA dynamical Lie algebras depend on the Hamiltonian, while sufficiently deep QAOA exhibits BPs for most maximum-cut graphs.
VI. STRATEGIES TO AVOID OR MITIGATE BPS
BP mitigation strategies reduce expressiveness, restrict dynamics, or operate in polynomial-sized spaces, but their effectiveness depends on measurements, state alignment, and reachability. Small dynamical Lie algebras can avoid some BPs, although such architectures are rare and may be classically simulable.
- Some mitigation methods encode relevant dynamics in polynomially rather than exponentially large subspaces of the operator space.
- Shallow circuits: Shallow circuits can reduce probabilistic and noise-induced BPs for certain measurements by limiting the number of gates.
- When the explored subspace has dimension O(poly(n)), loss variance can decay polynomially as Ω(1/dim(B_λ)), provided the input state and measurement operator align with it.
- Shallow circuits: Non-local measurements can still produce BPs in shallow circuits, and limited depth can create a reachability deficit or spurious local minima.
- Small dynamical Lie algebras: Circuits with polynomially growing dynamical Lie algebras can avoid probabilistic expressiveness-induced BPs when the initial state and measurement operator generate small modules.
- Small dynamical Lie algebras: Parametrized matchgate circuits have dim(g) ∈ O(n^2) and can avoid BPs, but polynomial dynamical Lie algebra dimension may also imply classical simulability.
- Small dynamical Lie algebras: Polynomial-sized dynamical Lie algebras are rare, while most circuits exhibit exponentially large algebras.
C. Variable structure ansatzes
Variable-structure ansatzes adapt circuit architecture to retain useful gradients while limiting problematic regions of highly expressive spaces. Initialization strategies can help heuristically, but optimization and noise mitigation do not generally remove the resource burden of BPs.
- Variable-structure PQCs: Variable-structure PQCs iteratively add or remove gates to lower the loss and preserve large gradients.
- Variable-structure PQCs: These methods use highly expressive gate sets while exploring only selected regions of the ansatz-architecture space.
- Initialization strategies: Variance-based BP analysis describes average landscape concentration and does not determine whether a particular initialization reaches a useful region.
- Initialization strategies: Proposed initialization strategies include restricted small-angle initialization, classical or quantum pre-training, parameter transfer, and iterative learning.
- Initialization strategies: Smart initialization has shown heuristic success, but it may lead to regions with large gradients that are poorly connected to a global minimum or are classically simulable.
- Strategies that cannot avoid BPs: Changing optimization methods still requires exponentially many measurements to identify a loss-minimizing direction in BP landscapes.
- Strategies that cannot avoid BPs: In worst cases, mitigating noise-induced BPs requires exponentially many resources, even for sub-logarithmic-depth circuits.
VIII. EXPONENTIAL CONCENTRATION ELSEWHERE
Barren Plateaus and exponential concentration appear across quantum generative models, kernel methods, quantum optimal control, tensor networks, and learnability theory. These connections reveal shared trainability and dimensionality issues while also identifying settings with guarantees or BP-avoidance strategies.
- Quantum generative models: Quantum Boltzmann machines have convex landscapes, allowing convergence guarantees with polynomial sample complexity.This provides a generative-model setting where the usual BP-related sampling difficulty is avoided.
- Kernel-based quantum algorithms: Kernel methods guarantee an optimal trained model for a fixed kernel and dataset despite exponential concentration in their feature-space estimates.Their guarantee concerns minimizing empirical loss for the chosen kernel and dataset.
- Quantum optimal control: Tools from quantum optimal control, including dynamical Lie algebra analysis, can diagnose BPs in parametrized quantum circuits.The review describes an intrinsic connection between exponentially concentrated control landscapes and BPs.
- Trainability of tensor networks: Tensor-network models exhibit BPs for global but not local losses, while isometric tensor-network states can avoid BPs.Related results also cover matrix product states, tree tensor networks, and multiscale entanglement renormalization ansätze.
- Exponential concentration elsewhere: BP-like exponential concentration has been studied across quantum generative models, kernel-based algorithms, quantum optimal control, tensor networks, and learning theory.The review connects BP research to multiple neighboring areas of quantum computing and machine learning.
- Learnability: BP results and statistical-query no-go theorems both link learning difficulty to a curse of dimensionality, although BPs remain algorithm-dependent.Statistical-query tools can provide algorithm-independent hardness bounds for specific function classes.
- Beyond barren plateaus: BPs do not fully resolve whether variational quantum computing can achieve quantum advantage, because local minima, solution quality, and classical simulability remain additional challenges.The review highlights exponentially many sub-optimal local minima and the possibility that removing them may require exponentially deep circuits.
X. LINK WITH VANISHING GRADIENT PROBLEM
The review distinguishes quantum BPs from classical vanishing gradients, while examining shared precision and trainability concerns. It argues that progress depends on parameterizations biased toward relevant problems and on jointly analyzing BPs with local minima, expressivity, and simulability.
- Differences between the two phenomena: Quantum BPs arise as the number of qubits increases, whereas the best-known classical vanishing-gradient problem arises with circuit or network depth.The review characterizes the quantum issue as a curse of dimension and the classical issue as a depth-related weight-norm problem.
- Classical solutions to vanishing gradients: Classical remedies such as normalization and ReLU activations target depth-induced shrinking weights and therefore do not appear directly applicable to quantum BPs.The review notes that these techniques can still fail in some classical situations.
- Classical solutions to vanishing gradients: Restricted-angle initialization and pre-training can enlarge quantum gradients, but under-parameterized quantum landscapes may contain numerous local minima.This makes direct transfer of classical gradient remedies less promising in quantum settings.
- The role of exponentially large spaces: The relevant issue is how exponentially large spaces are used, not merely their size; useful parameterizations should be naturally biased toward problems of interest.Classical models can also explore exponentially large spaces, so the review emphasizes the structure of the path through those spaces.
- The cost of precision: Quantum expected-value estimates require stochastic sampling whose cost scales as 1/ϵ^2 or, with clever methods, 1/ϵ, unlike classical precision scaling often given by O(log(1/ϵ)).The review notes that some classical models also use stochastic sampling and suggests deterministic readouts as a possible quantum strategy.
- Implications and outlook: BP causes and BP-avoiding architectures are increasingly understood, but their ultimate impact on variational-model scalability remains unresolved.The review attributes the primary causes to a curse of dimensionality while leaving scalability as an open question.
- Implications and outlook: Because BPs are average-case notions, exponentially narrow regions may still contain substantial gradients whose accessibility and usefulness for training remain open questions.The review calls for a holistic analysis combining BPs with local minima, expressivity limitations, and classical simulability.
- Implications and outlook: BP research identifies settings where quantum advantage is unlikely and may guide the search for practical variational advantage and methods beyond standard procedures.The review frames this as an ongoing role for BP studies rather than a completed determination of quantum advantage.