Source-linked AI summary

A Lie Algebraic Theory of Barren Plateaus for Deep Parameterized Quantum Circuits

Michael Ragone, Bojko N. Bakalov, Frédéric Sauvage, Alexander F. Kemper, Carlos Ortiz Marrero, Martin Larocca, M. Cerezo

arXiv:2309.09342v3quant-ph

TL;DR

Barren plateaus hinder training because loss landscapes can concentrate exponentially, while their known sources have often been studied separately. The paper develops a Lie algebraic theory for sufficiently deep parametrized quantum circuits, deriving exact variance expressions under stated state, observable, and noise conditions. The resulting framework unifies circuit expressiveness, entanglement, locality, and noise through the circuit’s dynamical Lie algebra.

  • Problem

    Barren plateaus hinder variational quantum computing, while circuit expressiveness, input entanglement, observable locality, and hardware noise have lacked a single unifying theory.

  • Method

    The paper studies the dynamical Lie algebra generated by a deep parametrized circuit and uses component-wise design assumptions to calculate loss variance.

  • Results

    The theory gives exact loss-variance formulas, including conditions under which the mean and variance receive contributions from the center or simple components of the dynamical Lie algebra.

  • Takeaways & Limitations

    The framework places known barren-plateau sources under one Lie algebraic description and relates exponential concentration to dynamical-Lie-algebra dimension and generalized purities.

Abstract

from arXiv · show

Variational quantum computing schemes train a loss function by sending an initial state through a parametrized quantum circuit, and measuring the expectation value of some operator. Despite their promise, the trainability of these algorithms is hindered by barren plateaus (BPs) induced by the expressiveness of the circuit, the entanglement of the input data, the locality of the observable, or the presence of noise. Up to this point, these sources of BPs have been regarded as independent. In this work, we present a general Lie algebraic theory that provides an exact expression for the variance of the loss function of sufficiently deep parametrized quantum circuits, even in the presence of certain noise models. Our results allow us to understand under one framework all aforementioned sources of BPs. This theoretical leap resolves a standing conjecture about a connection between loss concentration and the dimension of the Lie algebra of the circuit's generators.

I. INTRODUCTION

Barren plateaus make generic parametrized quantum circuits difficult to train because losses and gradients concentrate exponentially. The paper develops a Lie algebraic framework that unifies BP sources and computes loss variance for sufficiently deep circuits.

  • Motivation: Barren plateaus exponentially concentrate the loss and its gradients as problem size grows, requiring exponentially many measurement shots for trainability.This concentration prevents precise identification of loss-minimizing directions.
  • Motivation: Known BP sources include circuit expressiveness, measurement-operator locality, initial-state entanglement, and hardware noise.Prior studies generally analyze these sources in restricted scenarios, limiting their interconnection.
  • Contribution: The paper presents a general Lie algebraic theory applicable to any deep, unitary, parametrized quantum circuit architecture.The framework studies the Lie group and dynamical Lie algebra generated by the circuit.
  • Contribution: The theory computes loss variance exactly when the observable or input state belongs to the complexified dynamical Lie algebra, including specified noise models.It incorporates SPAM and coherent noise and generalizes locality, entanglement, expressiveness, and hardware noise within one framework.
  • Definitions: A barren plateau is defined by loss variance scaling as O(1/b^n) for some b > 1.The paper focuses on loss concentration because it implies concentration of partial derivatives.
  • Dynamical Lie algebra: The dynamical Lie algebra is the commutator-closed span of circuit generators and quantifies the circuit’s ultimate expressiveness.Its reductive decomposition consists of commuting simple ideals and an abelian center.
  • Dynamical Lie algebra: For sufficiently deep circuits forming component-wise 2-designs, parameter averages can be replaced by invariant Haar integrals evaluated with Lie-algebraic tools.The paper also bounds layers needed for approximate designs and the variance deviation from exact designs.
  • Dynamical Lie algebra: The framework uses generalized g-purity, defined from projection onto the complexification of an operator subalgebra and an orthonormal basis under the Hilbert–Schmidt inner product.This quantity captures how operators align with the circuit’s dynamical Lie algebra.

C. Main result

The main theorem gives an exact variance-based account of barren plateaus through Lie-algebra components, state and operator g-purities, and circuit expressiveness. It shows how exponential concentration arises from any of three sources and how these sources combine across components.

  • Theorem 1: Theorem 1 gives an exact loss-variance formula when the observable or input state lies in the Lie algebra's associative algebra.The loss mean vanishes on semisimple components, while the center contributes to the mean; conversely, variance vanishes for the center and retains simple-component contributions.
  • Theorem 1: The loss mean depends only on projections onto the DLA center, and an abelian DLA can produce a completely flat landscape when the state or observable commutes with it.For a centerless DLA, the mean loss is zero.
  • Three BP sources: Each variance term depends on the component dimension and the corresponding g-purities of the input state and measurement operator.The g-purity of a state measures generalized entanglement, while operator g-purity characterizes generalized locality.
  • Three BP sources: If dim(g), 1/Pg(ρ), or 1/Pg(O) grows as Ω(b^n) for b > 2, the loss has a barren plateau.These are the three sources identified by the theorem: circuit expressiveness, generalized entanglement of the input state, and generalized nonlocality of the measurement operator.
  • Ansatz expressiveness: An exponentially large DLA causes exponential loss concentration regardless of the initial state or measurement operator, whereas polynomial DLA dimension does not induce a BP by itself.The variance is inversely proportional to dim(g), establishing the conjectured connection between loss concentration and circuit expressiveness.
  • Measurement locality: Generalized-local operators maximize variance, while highly generalized-nonlocal measurements can produce a BP regardless of DLA dimension.For nonsimple DLAs, exponential concentration requires the relevant components to exhibit sufficiently large dimensions, generalized-entangled states, or generalized-nonlocal measurements.

E. Incorporating the effects of noise

The Lie algebraic framework interprets noise-induced loss concentration through changes in generalized entanglement, measurement locality, or circuit expressiveness, while identifying boundaries of its applicability.

  • Noise mechanisms: State-preparation noise reduces variance when it decreases the state’s g-purity, as occurs under global depolarizing noise.The noise maps P_gj(ρ) to a factor scaled by (1−p)^2.
  • Unified interpretation: Noise-induced concentration is unified as generalized entanglement in the state, reduced generalized locality of the observable, or increased circuit expressiveness.These mechanisms extend standard intuitions about entanglement, locality, and expressiveness into the DLA framework.
  • Noise mechanisms: Measurement errors induce loss concentration when the noisy observable loses generalized locality, meaning its g-purity decreases.The inverse noise channel used to model measurement errors need not be a physical quantum channel.
  • Noise mechanisms: Coherent errors add generators to the circuit, potentially enlarging the DLA and increasing expressiveness while decreasing the loss variance.The exact variance change depends on the reductive decomposition of the enlarged DLA.
  • Scope and limitations: The theory applies under specific assumptions, including sufficiently deep group-forming circuits and ρ or O belonging to the DLA-related space.It does not directly cover interleaved realistic noise when the noisy circuit no longer forms a group, nor circuits lacking approximate 2-design behavior.
  • DLA structure: The commutant and its center determine invariant Hilbert-space subspaces and illuminate the DLA’s decomposition into commuting ideals.This decomposition corresponds to a basis that block-diagonalizes the matrices in the DLA.
  • DLA structure: Universal unstructured circuits tend toward controllable circuits with g = su(2^n), whereas circuits with symmetries have richer DLA structure.The paper notes that improved DLA computation would increase the practical importance of these results.

B. Examples of barren plateaus and their origins

The examples show how expressiveness, measurement locality, initial-state entanglement, and circuit depth produce or avoid barren plateaus within the Lie algebraic framework.

  • Expressiveness: A 2-design over SU(2^n) has g = su(2^n) and dim(g) = 4^n − 1, yielding zero mean loss and barren-plateau concentration.Because the Lie algebra contains all nonidentity Pauli operators, the result applies regardless of O and ρ.
  • Measurement locality: For product single-qubit circuits, a global observable such as X^⊗n produces Varθ[ℓθ(ρ,O)] ∈ O(1/2^n), while a local observable X1 in the DLA does not.Here the DLA is g = su(2)^⊕n, so measurement locality determines whether concentration occurs.
  • Measurement locality: For a local operator acting on qubit j, the variance depends on the purity of the reduced input state on qubit j.The example gives a value of 1/3 in the relevant variance expression.
  • Initial-state entanglement: Initial states with volume-law entanglement can induce a barren plateau even for inexpressive single-qubit circuits and local measurements.The reduced state on each qubit is exponentially close to maximally mixed in this setting.
  • Deep circuits and 2-designs: The 2-design analysis estimates variance using only the second moment, although the authors prove the design statement for any moment order t.The circuit is factorized into layers and represented through moment operators comparing its ensemble with the Haar ensemble over G.
  • Deep circuits and 2-designs: Deep circuits form approximate G 2-designs after enough layers, extending prior full-algebra results to arbitrary dynamical Lie algebras.Theorem 2 also allows the required layer count to be computed from the expressiveness of one layer.
  • Deep circuits and 2-designs: The single-layer expressiveness depends on gate distribution and parameter sampling, whose effects are left for future work.This limits how directly the layer-count result can be specialized across ansatz choices.

D. Variance for a circuit that does not form a 2-design over G

The paper bounds how finite-depth circuits differ from the exact 2-design variance and shows exponential convergence with depth. It also establishes that loss variance decomposes across simple or abelian Lie-algebra components.

  • D. Variance for a circuit that does not form a 2-design over G: The section compares the variance of an L-layer circuit ensemble with the variance obtained when the circuit forms a 2-design over G.The exact comparison is formulated using the ensemble EL and moment superoperator MEL.
  • D. Variance for a circuit that does not form a 2-design over G: Theorem 3 bounds the difference between finite-depth variance and the 2-design variance when either ρ or O belongs to ig.The theorem applies to any dynamical Lie algebra g contained in u(2^n).
  • D. Variance for a circuit that does not form a 2-design over G: The variance converges exponentially fast in L to the 2-design variance, with the rate determined by the expressiveness of a single layer.Thus, in the deep-circuit regime, the 2-design assumption rapidly becomes exact.
  • E. Variance as a sum over simple or abelian components: The compact Lie group G decomposes into a product of component groups Gj associated with the direct-sum decomposition of the dynamical Lie algebra.A circuit unitary can therefore be represented as a tuple of component unitaries U = (U1, ..., Uk).
  • E. Variance as a sum over simple or abelian components: When O belongs to ig, it decomposes into operators Oj in the corresponding Lie-algebra components, and the loss becomes a sum of component losses.The component operators and unitaries are paired within the same Gj.
  • E. Variance as a sum over simple or abelian components: The expectation and variance of the loss split into sums over the simple or abelian components of g.Proposition 1 formalizes this decomposition for any density matrix ρ and O = Σ Oj.
  • E. Variance as a sum over simple or abelian components: The component decomposition is enabled by treating each component loss as depending only on its own coordinates, with independent component variables.This makes the expectation and variance of the total loss additive across components.

F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra

This section specializes the variance theory to states or observables diagonalizable with a Cartan subalgebra. In simple Lie algebras, the variance is controlled by the corresponding weight, and is largest for a highest-weight state.

  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: The special case assumes that ρ or O is simultaneously diagonalizable with a Cartan subalgebra h of g.For ρ, this means [ρ, h] = 0.
  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: For su(2), spin-S basis states satisfy Sz|m⟩ = m|m⟩, and the variance for O = Sz/S is m^2/(3S^2).Thus concentration depends on m, identified with the initial state’s g-purity or generalized entanglement.
  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: A weight basis diagonalizes the Cartan subalgebra, assigning each basis vector a weight λv.A pure weight state is a projector onto one such basis vector, while a weight state is diagonal in the weight basis.
  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: For a simple Lie algebra and a Cartan-diagonal density matrix, Corollary 2 gives the loss variance in terms of the associated weight.The result also holds with the roles of ρ and O exchanged.
  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: The variance depends on the norm of the weight, not on the particular Cartan subalgebra chosen.This follows because all Cartan subalgebras are conjugate.
  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: Among weight states in an irreducible representation, the highest-weight state maximizes the variance.Its weight has the largest norm among the representation’s weights.
  • F. Special case: ρ or O simultaneously diagonalizable with a Cartan subalgebra: The derivation estimates the variance using a trace bilinear form normalized through simple coroots and the highest weight’s integer evaluations.The resulting bound is obtained by substituting this estimate into the variance formula.

SUPPLEMENTAL INFORMATION FOR “A LIE ALGEBRAIC THEORY OF BARREN PLATEAUS FOR DEEP PARAMETERIZED QUANTUM CIRCUITS”

The supplemental information provides additional details and complete proofs of the paper’s main results.

  • The supplemental information contains additional details and complete proofs of the main results.

V. PROOF OF PROPOSITION 1

The proof establishes component-wise loss additivity and derives Theorem 1 using Haar moment operators as orthogonal projections onto group invariants. This yields distinct mean and variance contributions from abelian and simple components.

  • V. PROOF OF PROPOSITION 1: For two components, the loss separates into terms depending on θ1 and θ2 because operators in one component commute with the other component’s group.The general k-component case follows analogously.
  • V. PROOF OF PROPOSITION 1: The loss is expressed as a sum of independent component random variables, so its expectation and variance equal sums of the component expectations and variances.The product Haar measure ensures each component term can be averaged over its own coordinates.
  • VI. PROOF OF THEOREM 1 AND OF EQ. (28): Theorem 1 states that the mean vanishes on the semisimple component and retains only abelian contributions.The proof uses the first moment operator and the triviality of the center for simple components.
  • VI. PROOF OF THEOREM 1 AND OF EQ. (28): The variance vanishes for the abelian center and retains only simple-component contributions.This complements the mean decomposition and follows from the corresponding invariant-subspace structure.
  • V. PROOF OF PROPOSITION 1: The proof decomposes the dynamical Lie algebra and its group into simple or abelian components, reducing the argument to those cases.The loss, expectation, and variance are handled component by component.
  • VI. PROOF OF THEOREM 1 AND OF EQ. (28): The t-moment operator is an orthogonal projection onto operators invariant under conjugation by G.Self-adjointness and Haar-measure invariance establish the projection property.
  • VI. PROOF OF THEOREM 1 AND OF EQ. (28): For a simple Lie algebra, the invariant subspace of g ⊗ g is one-dimensional and spanned by the split quadratic Casimir element.This follows by identifying invariant tensors with equivariant endomorphisms and applying Schur’s lemma.
  • VI. PROOF OF THEOREM 1 AND OF EQ. (28): The proof computes the variance by projecting O⊗2 onto the invariant Casimir direction and evaluating the resulting Hilbert–Schmidt inner products.The g-purity enters through the orthogonal projection of a Hermitian operator into the complexified Lie algebra.

VII. PROOF OF COROLLARY 1

The proof establishes that exponential growth in the dynamical Lie algebra dimension, or inverse generalized purities of the state or observable, forces exponentially small loss variance.

  • Corollary 1: dim(g) in Ω(b^n) with b > 2 implies a barren plateau, regardless of the initial state or measurement operator.The proof separately analyzes exponential dimension of the dynamical Lie algebra and exponential inverse purities of the state or observable.
  • Corollary 1: 1/P_g(O) in Ω(b^n) with b > 2 likewise implies a barren plateau.This is the corresponding observable-purity case in the corollary.
  • Corollary 1: 1/P_g(ρ) in Ω(b^n) with b > 2 implies exponentially concentrated loss variance, irrespective of the dynamical Lie algebra dimension.The argument bounds the generalized purity contribution from the initial state.

VIII. PROOF OF THEOREM 2

The proof shows that sufficiently deep layered circuit ensembles converge toward the Haar moments on the generated group without assuming an approximate 2-design at the outset.

  • Approximate-design convergence: Theorem 2 states that an L-layered circuit forms an ϵ-approximate G t-design once L satisfies the theorem’s depth condition.The theorem applies to general t and does not begin by assuming that the circuit already forms an approximate 2-design.
  • Scope boundary: The required depth need not be uniform in the moment order t.The authors explicitly identify this as a scope limitation of the depth choice.
  • Moment-operator decomposition: The proof decomposes the moment operator into the invariant subspace and its orthogonal complement.The ensemble acts as the identity on G-invariants, while the complement is treated through the nontrivial eigenvalues.
  • Single-layer contraction: All eigenvalues of the single-layer moment operator on the non-invariant complement satisfy |λ| < 1.This is established in Lemma 6 and is the key contraction property used for deeper circuits.
  • Depth amplification: Composing L layers raises the largest nontrivial singular value to |λ|^L, producing exponential suppression with depth.Self-adjointness identifies singular values with absolute eigenvalues, and layer composition yields the Lth power.

A. Example of a use-case of Theorem 2

The use-case applies the moment-operator framework to hardware-efficient circuits by representing single-layer moments compactly and analyzing their dominant eigenvalue.

  • Moment-operator construction: The single-layer moment operator can be computed by decomposing each layer into independently sampled gate ensembles, including subgroup or U(1) representations.The construction uses commutant bases and Weingarten matrices for the component ensembles.
  • Circuit architecture: The example considers alternating neighboring two-qubit layers whose gates are independently sampled from local 2-designs over SU(4).The circuit uses a brick-like hardware-efficient architecture.
  • Eigenvalue behavior: For n = 2, the largest eigenvalue λ_max of A^(2)_E1 is 0, while for larger systems it appears to saturate near 0.639.The reported calculations cover n = 2, ..., 90.
  • Scope boundary: The example specializes to local SU(4) gates, although the method can extend to architectures whose gates are sampled from subgroups.Exploring such generalized cases is left for future work.

IX. PROOF OF THEOREM 3

Theorem 3 bounds the variance difference between a finite-depth circuit ensemble and a 2-design over its generated group using moment-operator norms.

  • Variance bound: Theorem 3 bounds the difference between the loss variance for an L-layer ensemble E_L and that of a 2-design over G.The result assumes a dynamical Lie algebra g contained in u(2^n) and requires either the state ρ or observable O to lie in i g.
  • Proof strategy: The proof expresses second moments through ensemble moment operators and controls the variance difference with Schatten norm inequalities.It separates the second-moment contribution from the difference of squared means.
  • Lower-order moments: An ϵ-approximate G t-design also forms an ϵ-approximate s-design for every 1 ≤ s ≤ t.Lemma 8 supplies the lower-order moment control used in the theorem’s analysis.
  • Higher moments: The framework extends beyond variance because higher-moment convergence can be bounded by adapting the same moment-operator argument.The paper notes that this also supports discussion of quantities such as the tth cumulant.

X. NUMERICAL SIMULATIONS

Numerical simulations test the Lie-algebraic variance formula across four circuit settings and find close agreement between theory and estimates. They also show that generalized input-state entanglement, rather than globality or standard entanglement alone, can determine barren-plateau behavior.

  • Simulation setup: 5,000 random circuit initializations per system size were used to estimate cost variances as circuit depth increased toward an approximate design.System sizes ranged from n ∈ [3, 15].
  • Simulation setup: Theoretical variances from Eq. (9) match numerical estimates closely across all four setups.The remaining discrepancies are attributed to sample-variance uncertainty.
  • Results: A global measurement operator does not necessarily produce a barren plateau when it belongs to ig.Setup 1 uses the global observable O = X1Yn with the highest-weight input state.
  • Results: Single-qubit rotations slightly reduce variance without inducing a barren plateau, because the variance remains polynomially vanishing.This setting decreases g-purity while leaving standard purity and standard entanglement unchanged.
  • Results: Applying random two-qubit rotations generates enough generalized entanglement to induce a barren plateau, evidenced by exponentially decaying variance.This behavior contrasts with the single-qubit-rotation setup.
Loading 2309.09342v3…