Source-linked AI summary

Diagnosing Barren Plateaus with Tools from Quantum Optimal Control

Martin Larocca, Piotr Czarnik, Kunal Sharma, Gopikrishnan Muraleedharan, Patrick J. Coles, M. Cerezo

arXiv:2105.14377v3quant-ph

TL;DR

The paper addresses limited understanding of gradient scaling in problem-inspired VQAs and whether they avoid barren plateaus. It uses quantum optimal control and dynamical Lie algebras to diagnose trainability, showing that controllability, Lie-algebra dimension, and input-state support determine gradient scaling. The results also establish limitations for controllable systems and certain near-maximally mixed inputs.

  • Problem

    The gradient scaling and barren-plateau behavior of problem-inspired ansatzes remain insufficiently understood despite their proposed trainability advantages.

  • Method

    The paper applies quantum optimal control tools to periodic ansatzes and analyzes the dynamical Lie algebra generated by their operators.

  • Results

    The analysis links barren-plateau behavior to controllability and Lie-algebra dimension, with input-state support determining trainability in subspace-controllable systems.

  • Takeaways & Limitations

    Trainability-aware ansatz design can use DLA analysis and input-state choice without requiring extra quantum resources.

  • Takeaways & Limitations

    Polynomially growing DLAs do not preclude barren plateaus when the input state is exponentially close to maximally mixed on its supported subspace.

Abstract

from arXiv · show

Variational Quantum Algorithms (VQAs) have received considerable attention due to their potential for achieving near-term quantum advantage. However, more work is needed to understand their scalability. One known scaling result for VQAs is barren plateaus, where certain circumstances lead to exponentially vanishing gradients. It is common folklore that problem-inspired ansatzes avoid barren plateaus, but in fact, very little is known about their gradient scaling. In this work we employ tools from quantum optimal control to develop a framework that can diagnose the presence or absence of barren plateaus for problem-inspired ansatzes. Such ansatzes include the Quantum Alternating Operator Ansatz (QAOA), the Hamiltonian Variational Ansatz (HVA), and others. With our framework, we prove that avoiding barren plateaus for these ansatzes is not always guaranteed. Specifically, we show that the gradient scaling of the VQA depends on the degree of controllability of the system, and hence can be diagnosed through the dynamical Lie algebra $\mathfrak{g}$ obtained from the generators of the ansatz. We analyze the existence of barren plateaus in QAOA and HVA ansatzes, and we highlight the role of the input state, as different initial states can lead to the presence or absence of barren plateaus. Taken together, our results provide a framework for trainability-aware ansatz design strategies that do not come at the cost of extra quantum resources. Moreover, we prove no-go results for obtaining ground states with variational ansatzes for controllable system such as spin glasses. Our work establishes a link between the existence of barren plateaus and the scaling of the dimension of $\mathfrak{g}$.

1 INTRODUCTION

VQAs offer a near-term route toward quantum advantage, but barren plateaus can make optimization untrainable as system size grows. This work uses QOC tools and dynamical Lie algebras to diagnose barren plateaus in periodic problem-inspired ansatzes.

  • Motivation: VQAs encode tasks in efficiently computable parametrized cost functions whose parameters are optimized classically on noisy quantum computers.They have been proposed for linear systems and dynamical quantum simulations.
  • Motivation: Barren plateaus cause cost-function gradients to vanish exponentially with system size, making optimization untrainable and requiring exponentially high precision.They are recognized as a major obstacle to scalable VQAs.
  • Motivation: Problem-inspired ansatzes are intended to constrain optimization toward spaces containing exact or approximate solutions, but their immunity to barren plateaus is not established.The paper challenges the common expectation that encoding problem structure guarantees trainability.
  • Contribution: QOC tools provide a common framework for diagnosing barren plateaus in periodic ansatzes including QAOA and HVA.The same procedure also applies to adaptive QAOA, quantum optimal control ansatzes, and quantum neural network architectures.
  • Contribution: The proposed diagnosis studies controllability through the dynamical Lie algebra generated by the ansatz operators.The paper organizes its results around different controllability scenarios.

2 VARIATIONAL QUANTUM ALGORITHMS

The paper formulates periodic variational circuits, defines barren plateaus through exponentially vanishing gradient variance, and relates trainability to ansatz expressibility. Highly expressive circuits approach 2-design behavior, where gradients become small.

  • General framework: The optimization minimizes a cost function determined by an input state, parametrized circuit, and Hermitian task operator.The input state acts on n qubits in a d-dimensional Hilbert space with d = 2^n.
  • General framework: Periodic Structure Ansatzes consist of layered parametrized circuits generated by Hermitian traceless operators, with some parameters optionally fixed.The layer parameters collectively define the circuit controls.
  • Barren plateaus: A barren plateau is defined by cost-function partial-derivative variance that vanishes exponentially with system size.In this regime, gradients are exponentially suppressed across the optimization landscape.
  • Expressibility: Ansatz expressibility compares the generated unitary distribution with the Haar distribution using deviations of second moments.The paper focuses on second-moment expressibility.
  • Expressibility: Greater expressibility is associated with smaller gradients, and the 2-design limit corresponds to a barren plateau.The expressibility superoperator norm decreases as the ansatz becomes more expressive.

3 QUANTUM OPTIMAL CONTROL

Quantum Optimal Control studies systematic manipulation of quantum dynamics through controllable generators and time-dependent controls. Its dynamical Lie algebra determines the reachable unitaries and states, providing the controllability framework used by the paper.

  • QOC framework: QOC manipulates quantum dynamical systems by varying time-dependent control fields alongside a fixed drift Hamiltonian and control Hamiltonians.The resulting propagator solves the Schrödinger equation.
  • QOC framework: Periodic Trotterized QOC propagators form Periodic Structure Ansatzes, linking pulse-level control to circuit-level VQAs.This establishes the shared variational structure used in the paper.
  • Dynamical Lie algebra: The generator set consists of the Hermitian traceless operators used to generate the ansatz unitaries.The DLA is obtained from these generators rather than from their individual elements alone.
  • Dynamical Lie algebra: The dynamical Lie algebra is the Lie closure of the generators under repeated nested commutators.It is a subalgebra of su(d), the traceless skew-Hermitian matrix algebra.
  • Reachability: The DLA determines the dynamical Lie group and therefore the unitaries and states reachable from an input state.This makes DLA analysis a route to studying control-system expressibility.
  • Controllability: Full-rank DLAs yield controllability, whereas proper subalgebras restrict reachable states; symmetries can partition the Hilbert space into invariant subspaces.A system may nevertheless be controllable within individual invariant subspaces.

4 MAIN RESULTS

The paper links barren-plateau behavior in periodic problem-inspired ansatzes to controllability, dynamical Lie algebra dimension, expressibility, and input-state support. Controllable systems can develop barren plateaus at sufficient depth, while subspace controllability makes gradient scaling depend on the relevant invariant subspace.

  • 4.1 Controllable systems: A controllable system forms an exponentially accurate approximate 2-design at an appropriate depth and consequently exhibits a barren plateau.The required approximation has ε ∈ O(1/2^n).
  • 4.1 Controllable systems: The layered Hardware Efficient Ansatz is controllable and can therefore exhibit barren plateaus, while the same result yields a no-go theorem for deep variational ansatzes targeting certain spin-glass ground states.The controllability proof for the Hardware Efficient Ansatz is presented as novel, and the spin-glass result concerns determining ground-state energies.
  • 4.2 Subspace controllable systems: For subspace-controllable systems, gradient-variance scaling is governed by the dimension of the invariant subspace containing the input state rather than the full Hilbert-space dimension.Thus, different input states can produce different trainability behavior within the same ansatz.
  • 4.2 Subspace controllable systems: In the XXZ example, inputs with a fixed or nonextensive excitation number may avoid barren plateaus, whereas the k = n/2 subspace has exponentially large dimension and exhibits one.The result follows from the invariant-subspace decomposition induced by the Hamiltonian symmetries.
  • 4.5 General case: linking gradient scaling to the dimension of the Lie algebra: Ansatz expressibility also suppresses gradients: highly expressible ansatzes in exponentially large subspaces can exhibit barren plateaus, and the input state remains crucial even when the DLA is small.A normalized cost removes variance growth caused solely by the increasing landscape scale, but exponentially mixed input states can still produce exponentially vanishing gradients.
  • 4.5 General case: linking gradient scaling to the dimension of the Lie algebra: The proposed general principle is that gradient-variance scaling is inversely related to dynamical Lie algebra dimension, enabling trainability diagnosis through theoretical DLA analysis.Polynomially growing sub-DLAs may yield only polynomially vanishing gradients, while exponentially growing sub-DLAs yield exponentially vanishing gradients; numerical DLA construction can scale poorly.

5 NUMERICAL SIMULATIONS

Numerical simulations across controllable, subspace controllable, subspace uncontrollable, and model-specific ansatzes verify that gradient-variance scaling tracks the dynamical Lie algebra dimension. The simulations also show that input states and small ansatz modifications can determine whether barren plateaus occur.

  • 5.1 Controllable systems: Controllable hardware-efficient systems exhibit barren plateaus, with gradient variance scaling as O(1/2^n) while dim(g) = 4^n − 1.The variance is polynomial in 1/dim(g) on a log-log scale, confirming Conjecture 1.
  • 5.2.1 The XXZ model: For the XXZ HVA, fixed excitation numbers m = 1, 2, …, 5 yield polynomially decreasing variance and no barren plateau, whereas m = n/2 yields exponential decay.Theoretical and numerical results agree, and barren plateaus are suggested whenever the invariant-subspace dimension scales as O(2^n).
  • 5.2.1 The XXZ model: The uncontrollable XXZ ansatz with m = n/2 still exhibits a barren plateau, although its variance is larger than in the corresponding controllable case.The results use the expressibility analysis because the controllable-system theorem does not apply.
  • 5.2.2 The Ising Model: Adding one longitudinal-field unitary changes the DLA dimension and changes trainability from polynomial variance decay in TFIM to exponential decay in LTFIM.The modification is parameterized by a single angle in each layer.
  • 5.2.2 The Ising Model: TFIM ansatzes show polynomially vanishing gradient variance and no barren plateau, while LTFIM ansatzes show exponential decay and a barren plateau under both boundary conditions.The corresponding DLA dimensions grow polynomially for TFIM and exponentially for LTFIM, verifying Conjecture 1.
  • 5.2.3 Erdös–Rényi model: For Erdös–Rényi graphs, the median gradient variance decays exponentially with system size, suggesting a barren plateau for a typical uniformly sampled graph.Across graphs, the variance remains linearly related to DLA dimension on a log-log scale.

6 DISCUSSION

The paper links barren-plateau behavior in periodic problem-inspired ansatzes to controllability, measured through the dynamical Lie algebra, and uses this link to guide ansatz design. Results also show that input states, ansatz modifications, and classical DLA analysis materially affect trainability assessment.

  • Main results: The framework diagnoses barren plateaus by analyzing the controllability of the system through the dimension of the ansatz-generated dynamical Lie algebra.The framework applies to periodic ansatzes including QAOA and HVA.
  • Main results: For controllable systems with full-rank DLA, the cost function exhibits a barren plateau because the ansatz converges to 2-designs.The paper also relates the depth needed for an ε-approximate 2-design to the expressibility of one layer.
  • Main results: In subspace-controllable systems, barren-plateau existence depends on the input state and gradient-variance scaling follows the dimension of the supported invariant subspace.Some input states can remain trainable while others exhibit barren plateaus.
  • Main results: Polynomially growing subspace DLAs can yield polynomially vanishing gradients, whereas exponentially growing DLAs should yield exponentially vanishing gradients.This conclusion is presented as a conjecture relating derivative-variance scaling to subspace-DLA dimension.
  • Numerical validation: Numerical simulations of hardware-efficient, QAOA, and HVA ansatzes matched the theoretical predictions across ground-state preparation and MAXCUT problems.The simulations covered XXZ, Ising, and Erdős–Rényi MAXCUT instances.
  • Implications: DLA-based trainability analysis can be performed classically, avoiding quantum-algorithm execution or quantum-computer access when testing ansatzes.The authors propose this as a resource-saving route toward trainability-aware ansatz design.
  • Implications: Adding an additional parametrized unitary per layer can change gradient scaling by changing controllability, so adaptive QAOA and quantum-control ansatz modifications require caution.Such additions can produce barren plateaus if the resulting system becomes controllable or subspace controllable.
  • Outlook: The framework is currently limited to certain periodic ansatz families, while extensions to non-periodic ansatzes and formal proof of the conjecture remain open.The paper also identifies broader quantum-optimal-control trainability analysis as future work.

Appendices

The appendices supply background on VQAs, ansatz families, barren plateaus, Haar integration, and dynamical Lie algebras. They also identify the principal ansatz classes covered by the framework and outline the computational difficulty of DLA construction.

  • Integration tools: Haar measure is finite, unique up to normalization, invariant under left and right unitary actions, and supports formulas for first and second moments.The appendices introduce these properties before applying symbolic integration identities.
  • Integration tools: Parameter-space integration can be related to integration over the unitary ensemble generated by parameter choices, and 2-design convergence permits replacement by Haar integration.This supports the symbolic moment calculations used later.
  • Ansatz families: The framework covers hardware-efficient, QAOA, Adaptive QAOA, HVA, and Quantum Optimal Control Ansatzes as periodic-structure ansatzes.The appendices distinguish problem-agnostic hardware-efficient circuits from problem-inspired constructions such as QAOA and HVA.
  • Barren plateaus: Barren plateaus are defined by gradients that are exponentially suppressed on average across the optimization landscape, making the cost function untrainable.The appendices present this as a central challenge for VQA success.
  • Dynamical Lie algebras: Direct DLA construction repeatedly adds nested commutators until a basis is obtained, with generic complexity polynomial in d=2^n and therefore exponential in qubit number n.Dense-matrix implementations can incur roughly O(d^2d^6) cost because linear independence checks use matrix decompositions.

F Proof of Theorem 1: Convergence of controllable systems to 2-designs

The proof establishes that a controllable periodic ansatz converges toward an approximate 2-design by analyzing its layered moment operator and its eigenstructure. Controllability ensures access to all unitaries up to phase.

  • Theorem 1: Theorem 1 states that a controllable periodic-structure ansatz forms an ε-approximate 2-design after sufficiently many layers.The proof measures convergence through the distance between the ansatz and Haar second-moment operators.
  • Expressibility: The single-layer expressibility is used to characterize the distance from the ansatz distribution to a 2-design.The proof explicitly denotes the expressibility of a single layer as the relevant quantity.
  • Proof strategy: The proof uses harmonic analysis to study convergence of the unitary distribution and compares the resulting second moment with the Haar moment operator.Haar invariance and symbolic integration identities supply the needed operator calculations.
  • Moment operators: For an L-layer ansatz, the moment operator is the L-th power of the single-layer moment operator.This layered structure is the basis for the convergence argument.
  • Controllability: Controllability implies that every unitary in SU(d) can be obtained for some parameter choice, forcing the relevant eigenvectors to coincide with those of the Haar second-moment operator.The Haar operator has two eigenvectors with eigenvalue 1 in the argument presented.
  • Convergence: The remaining moment-operator eigenvalues have modulus below 1, so repeated layers contract toward the two-design fixed subspace.The proof identifies the nontrivial eigenvalues as the convergence-limiting terms.

G Proof of Corollary 1: Rate of convergence of controllable systems to 2-designs

The corollary gives a layer-depth condition under which a controllable ansatz reaches an exponentially accurate approximate 2-design. The condition depends on the single-layer expressibility gap.

  • Corollary 1: If the single-layer expressibility gap satisfies δ(n)∈Ω(1/poly(n)), then L(n)∈Ω(n/δ(n)) layers suffice for the convergence guarantee.The condition assumes δ(n) vanishes no faster than polynomially with n.
  • Corollary 1: Under that depth condition, the ansatz is no worse than an ε(n)-approximate 2-design with ε(n)∈O(1/2^n).The n-dependence is included explicitly in both the layer count and approximation error.

H Proof of Proposition 1: Controllability leads to barren plateaus

For controllable systems, sufficient depth produces approximate 2-designs, causing exponentially vanishing gradient variance and barren plateaus.

  • Controllable systems can form ε-approximate 2-designs with ε ∈ O(1/2^n) at an appropriate depth scaling.This establishes the expressibility condition used in the proposition.
  • Therefore, controllable systems exhibit barren plateaus at the stated depth scaling.
  • The ε-approximate 2-design condition implies that the cost-function gradient variance vanishes exponentially with system size.The proof invokes prior results connecting ε ∈ O(1/2^n) to exponentially small derivative variance.

I Proof of Proposition 2: Controllability of the HEA and the Spin Glass model

The HEA and spin-glass generator sets generate full-rank dynamical Lie algebras, making both systems controllable through nested commutators.

  • The two generator sets in Proposition 2 generate full-rank dynamical Lie algebras and therefore controllable systems.
  • HEA: For the HEA, nested commutators generate all two-body terms, then higher-body operators, yielding all n-body Pauli operators.
  • Spin-glass model: For the spin-glass model, commutators generate single-qubit operators and the generator structure needed to recover the HEA generators.
  • HEA: The HEA dynamical Lie algebra is full rank, which implies that the hardware-efficient ansatz is controllable.
  • Spin-glass model: Assuming Gaussian-sampled coefficients are nonzero and distinct, a Vandermonde argument establishes linear independence of the relevant operators.
  • Spin-glass model: The spin-glass dynamical Lie algebra is full rank, implying that the spin-glass system is controllable.

J Proof of Theorem 2: Variance in subspace controllable systems

For a reducible system controllable on an invariant subspace, the gradient-variance analysis reduces to that subspace when the initial state lies entirely within it.

  • A symmetry operator decomposes the Hilbert space into invariant subspaces, and the dynamical Lie algebra correspondingly becomes block diagonal.
  • If the system is controllable on H_k and the initial state lies in H_k, the gradient variance is determined by reduced operators on that subspace.
  • The proof expands the ansatz around a parameter into unitaries before and after that parameter, then averages the resulting derivative expression.
  • Subspace controllability permits replacing the unitary distributions with ε-approximate 2-designs over U(d_k).
  • The derivation relies on the initial state occupying one invariant subspace; states spread across multiple subspaces require a more careful analysis.

K Proof of Corollary 2: Exponentially growing subspaces have barren plateaus

When a controllable invariant subspace grows exponentially with system size and fourth-moment traces satisfy the stated scaling, the cost exhibits a barren plateau.

  • If Tr[(H_μ)^4] and Tr[O^4] are O(2^n), a controllable subspace with d_k ∈ O(2^n) exhibits a barren plateau.
  • The proof bounds the gradient variance using Hilbert-Schmidt distances and positivity of the projected operator.
  • For exponentially growing d_k and the stated fourth-trace assumptions, the variance is upper bounded by a function that vanishes exponentially with n.
  • The fourth-trace assumption covers projectors of arbitrary rank and operators with polynomially many Pauli-string terms.

L Proof of Theorem 3: Expressibility in the subspace

Theorem 3 bounds gradient variance for a reducible system by analyzing expressibility within an invariant subspace. The proof rewrites the variance using subspace-restricted second moments and operator norms.

  • Theorem 3: Theorem 3 considers an initial state in an invariant subspace of dimension d_k and upper-bounds the variance of a cost-function derivative.The bound is formulated for reducible systems with ρ ∈ H_k.
  • Variance decomposition: The derivative variance is expressed in terms of commutators involving the Hamiltonian generator, observable, unitary distributions, and input state.The proof introduces X and Y from these commutator structures before applying norm-based bounds.
  • Expressibility: The proof uses the expressibility superoperator for second moments restricted to the k-th invariant subspace.The relevant expressibility is evaluated over the unitary distributions associated with the two circuit portions.
  • Bounding steps: Triangle inequality and Cauchy–Schwarz convert the variance expression into an upper bound involving Frobenius norms and the quantity defined in Theorem 2.The derivation applies these inequalities after substituting the subspace-restricted second-moment expressions.

M Proof of Proposition 3: Variance on the irreducible representations of SU(2)

The SU(2) analysis derives gradient-variance expressions for a toy ansatz under 2-design assumptions and validates them numerically. It shows that trainability depends on the chosen initial state and system structure.

  • Proposition 3: Proposition 3 derives the variance of a cost-function derivative when the circuit’s before-and-after unitary distributions converge to 2-designs on SU(2).The proposition considers a parameter associated with a circuit generator and gives the resulting variance expression.
  • Toy model: The toy ansatz uses alternating exponentials generated by S_x and S_y in the d-dimensional irreducible representation of su(2).The spin basis is chosen as eigenvectors of S_z, with dimension d determined by the spin representation.
  • Variance calculation: The variance calculation averages the squared derivative over SU(2) using Euler-angle parameterization, Haar integration, and a two-copy expression for the input state.The cost derivative is represented through a commutator with the observable, and the mean derivative vanishes because commutators are traceless.
  • Variance calculation: The derivation reduces the variance to three nonzero matrix-element contributions and then combines their integrated terms into the final expression.Selection rules and Clebsch–Gordan orthogonality eliminate the other terms before the final aggregation.
  • Numerical validation: For L = 100 and 1200 random initializations, the theoretical prediction matches simulations: |m⟩ = |1/2⟩ is trainable, whereas |m⟩ = |S⟩ exhibits a barren plateau.The comparison uses the normalized cost function involving S_x + S_y + S_z.
Loading 2105.14377v3…