Source-linked AI summary
Connecting ansatz expressibility to gradient magnitudes and barren plateaus
Zoë Holmes, Kunal Sharma, M. Cerezo, Patrick J. Coles
TL;DR
VQAs need ansätze that are expressive enough to access solutions yet trainable through sufficiently large gradients. The paper extends barren-plateau analysis to arbitrary ansätze by relating gradient variance to expressibility, finding that highly expressive ansätze generally have flatter landscapes, while numerics identify mitigation strategies.
Problem
VQA ansätze must balance access to solutions through expressibility with trainability through sufficiently large gradients.
Method
The paper analytically extends the 2-design barren-plateau result to arbitrary ansätze by bounding cost-gradient variance from their distance to a 2-design.
Results
Highly expressive ansätze have smaller cost-gradient variance and flatter landscapes, while numerical studies examine reduced depth, correlated parameters, and restricted rotations.
Takeaways & Limitations
Problem-inspired ansätze need not be highly expressive, and correlating parameters or restricting rotation angles may mitigate barren plateaus in the reported numerics.
Abstract
from arXiv · showhide
Parameterized quantum circuits serve as ansätze for solving variational problems and provide a flexible paradigm for programming near-term quantum computers. Ideally, such ansätze should be highly expressive so that a close approximation of the desired solution can be accessed. On the other hand, the ansatz must also have sufficiently large gradients to allow for training. Here, we derive a fundamental relationship between these two essential properties: expressibility and trainability. This is done by extending the well established barren plateau phenomenon, which holds for ansätze that form exact 2-designs, to arbitrary ansätze. Specifically, we calculate the variance in the cost gradient in terms of the expressibility of the ansatz, as measured by its distance from being a 2-design. Our resulting bounds indicate that highly expressive ansätze exhibit flatter cost landscapes and therefore will be harder to train. Furthermore, we provide numerics illustrating the effect of expressiblity on gradient scalings, and we discuss the implications for designing strategies to avoid barren plateaus.
I. Introduction
VQAs require ansätze that are both expressive enough to contain good solutions and trainable enough to provide usable cost gradients. The paper analytically connects these properties, showing that greater expressibility generally produces flatter landscapes while numerics examine mitigation strategies.
- Successful VQAs require ansätze that contain near-optimal solutions while maintaining sufficiently featured cost landscapes for parameter training.
- 2-design ansätze exhibit barren plateaus in which cost-gradient variance vanishes exponentially with qubit number.
- The paper extends barren-plateau analysis from exact 2-designs to arbitrary ansätze by bounding gradient variance using distance from a 2-design.
- Greater expressibility yields smaller gradient variance and flatter cost landscapes, although an ansatz need only contain the problem’s solution to succeed.
- Numerics tune circuit depth, parameter correlations, and rotation directions or angles to study how reduced expressibility affects gradient scaling.
A. General framework
The framework models VQAs as classical optimization over parameters of quantum circuits, with costs evaluated on quantum hardware. It emphasizes that ansätze must contain problem solutions, while locality and measurement considerations also shape cost-function design.
- A. General framework: VQAs evaluate a parameterized circuit’s cost or gradient on a quantum computer and use a classical optimizer to minimize it over circuit parameters.
- A. General framework: A faithful cost function must have its minimum correspond to the solution of the optimization problem.
- A. General framework: Generalized costs can combine multiple input states and measurement operators, supporting quantum machine-learning approaches using training data.
- A. General framework: Costs are classified as global when their measurement operator acts on all qubits and k-local when it acts nontrivially on at most k qubits.
- A. General framework: Parameterized circuits are built from fixed unitaries and rotations e^(-iθ_jV_j), with rotation angles typically treated as independent parameters.
B. Expressibility
Expressibility measures how uniformly an ansatz explores the unitary group and can be defined through its distance from Haar-random unitary moments. Problem-inspired ansätze may be complete yet inexpressive, whereas problem-agnostic ansätze require broader exploration for completeness.
- B. Expressibility: An ansatz is complete for a problem when its accessible unitaries overlap the problem’s solution-unitary space.
- B. Expressibility: Expressive ansätze explore the unitary space broadly and uniformly, increasing completeness when the locations of solution unitaries are unknown.
- B. Expressibility: Problem-inspired ansätze can be complete but inexpressive, while problem-agnostic ansätze need sufficient expressibility to guarantee completeness across problems.
- B. Expressibility: Expressibility compares the unitary ensemble generated by a circuit with the uniform Haar distribution over U(d).
- B. Expressibility: Matching Haar averages through the t-th moment means the ansatz forms a t-design; the analysis focuses on t = 2.
- B. Expressibility: The diamond norm provides an operationally meaningful, cost-independent distance for distinguishing quantum operations and quantifying expressibility.
C. Gradient Magnitudes
Containing a solution is not enough for VQA success: gradients must fluctuate sufficiently away from zero to guide optimization. The framework uses unbiased gradients, variance, and Chebyshev’s inequality to characterize trainability.
- C. Gradient Magnitudes: VQA training requires cost landscapes with gradients large enough to guide discovery of an ansatz-contained solution.
- C. Gradient Magnitudes: For parameter θ_k, the relevant gradient component is the partial derivative ∂_kC = ∂C_ρ,H(θ)/∂θ_k.
- C. Gradient Magnitudes: The average cost-gradient component over parameters is zero, so the gradients are unbiased rather than systematically directed.
- C. Gradient Magnitudes: Trainability depends on fluctuations around zero: Chebyshev’s inequality bounds the probability that a gradient deviates from its mean using its variance.
- C. Gradient Magnitudes: Small variance makes nonzero gradients unlikely and can require extremely precise measurements to identify a descent direction.
D. Barren Plateaus
Barren plateaus occur when gradient variance vanishes exponentially with system size, particularly when circuit ensembles form 2-designs.
- D. Barren Plateaus: A barren plateau is characterized by cost gradients that vanish exponentially with the number of qubits n.The probabilistic formulation follows when the gradient variance vanishes exponentially, via Chebyshev’s inequality.
- D. Barren Plateaus: For a layered ansatz, differentiating with respect to a parameter effectively splits the circuit into independent left and right portions.The independence follows from assuming uncorrelated parameters.
- D. Barren Plateaus: If either circuit portion forms a 2-design, the variance of the cost gradient vanishes exponentially with n.The relevant cases include the left portion, the right portion, or both portions forming 2-designs.
- D. Barren Plateaus: Maximally expressive ansätze therefore exhibit barren plateaus under the 2-design condition.The gradient-variance scaling is made explicit through an n-dependent factor and a prefactor determined by the state, Hamiltonian, and circuit.
A. Analytic Bounds
The paper derives generic upper bounds linking gradient variance to ansatz expressibility, showing that greater expressibility produces flatter landscapes while not guaranteeing trainability for inexpressive circuits.
- A. Analytic Bounds: The converse fails: inexpressive ansätze can also have vanishing gradients, including commuting rotations and tensor-product rotations under global costs.Thus expressibility cannot provide a meaningful lower bound on gradient magnitudes.
- A. Analytic Bounds: The main result upper-bounds cost-gradient variance for arbitrary layered ansätze using their distance from 2-designs as an expressibility measure.The bound combines the 2-design gradient variance with expressibility-dependent distances for the left and right circuit portions.
- A. Analytic Bounds: Higher expressibility, corresponding to smaller 2-design distance, yields a smaller upper bound on gradient variance and flatter, harder-to-train landscapes.The conclusion uses the unbiasedness of the cost gradient in addition to the variance bound.
- A. Analytic Bounds: The bounds apply to any generic ansatz, with relative tightness depending on the differentiated layer and whether the left, right, or both circuit portions are closer to 2-designs.Bounds associated with final, initial, and middle layers can be respectively more informative under these conditions.
- A. Analytic Bounds: The analytic results extend to costs with multiple input states and measurements, including settings using training data.This extension is stated for quantum machine-learning approaches that utilize training data.
- A. Analytic Bounds: If the expressibility correction scales non-exponentially, the upper bound allows gradient variance to remain non-vanishing, leaving room to avoid barren plateaus.When the correction vanishes or grows exponentially in the relevant way, gradient variance again vanishes exponentially.
- A. Analytic Bounds: The bounds also imply that increasing expressibility increases concentration of cost-function values around their mean.This follows from the established association between cost concentration and gradient variation.
- A. Analytic Bounds: For local costs, one bound becomes exponentially loose because ||H||2 scales exponentially with system size, motivating a diamond-norm reformulation.The alternative formulation avoids the same looseness for large local-cost systems, although another bound may still become loose when ||H||1 scales exponentially.
B. Numerical Simulations
The numerics examine how circuit depth, parameter correlations, rotation directions, and angle ranges affect gradient scalings as expressibility is varied. Reduced expressibility can increase gradients, but the effect depends on cost locality and initialization.
- Circuit depth: Reducing circuit depth preserves exponentially vanishing partial derivatives for global costs, while shallow circuits yield approximately constant scaling for local costs.For local costs, exponential decay appears at D ⪆100, whereas D ⪅50 gives approximately constant scaling for n ⪆8.
- Correlating parameters: Correlating both qubits and layers produces the least expressive ansatz and approximately constant gradient-variance scaling.Correlating only qubits or only layers reduces the system-size scaling, but exponential scaling remains.
- Restricting rotation direction: Restricting rotation directions removes exponential gradient scaling for local costs but not for global costs.Using only z rotations makes the cost landscape entirely flat because the ansatz commutes with the local Hamiltonian.
- Restricting rotation angle: Restricting angle ranges does not change partial derivatives under random initialization, because it limits the explored region rather than the cost landscape.The partial derivatives for different range factors r perfectly overlap with random initialization.
- Restricting rotation angle: When initialized near the solution, small angle ranges remove exponential scaling for local costs and weakly reduce it for global costs.For local costs this occurs around r ⪅0.1; for global costs the effect is visible only near r ≈0.025 in these data.
- Outlook for ansatz design: A practical design strategy is to begin with shallow, highly correlated circuits and gradually increase depth and decorrelate parameters during optimization.Angle-range restriction also appears useful, but requires initialization close to the solution and therefore prior knowledge or effective pre-training.
IV. Discussion
The discussion extends barren-plateau results from exact 2-designs to arbitrary ansätze by linking expressibility with gradient variance. The analysis and numerics identify correlations, caveats, and possible strategies for mitigating barren plateaus.
- Analytic extension: Theorems 1 and 2 extend the barren-plateau result from exact 2-design ansätze to arbitrary ansätze, including approximate 2-designs.The resulting bounds can provide gradient-variance estimates for realistic ansätze that are not exact 2-designs.
- Caveats: For highly correlated circuits, gradient variance can exceed the bounds derived for uncorrelated circuits, motivating generalized bounds that account for such correlations.The authors leave this generalization to future work.
- Expressibility–trainability trade-off: Increasing ansatz expressibility can result in smaller cost gradients, linking expressibility and trainability through the distance from a 2-design.The bounds relate gradient magnitudes to the ansatz’s expressibility.
- Numerical relationship: Numerics typically show a strong correlation between expressibility and gradient variance, especially for local cost functions, although the bounds are not perfectly tight.The lack of tightness may arise from repeated use of triangle and Cauchy–Schwarz inequalities in the derivation.
- Caveats: The numerical results are problem specific because they depend on the chosen cost function and ansatz, so the universality of the observed trends remains unresolved.Further work is needed to determine how broadly these trends apply and whether analytic support can be obtained.
- Mitigation strategies: The numerical trends suggest that correlating parameters and restricting rotation angles, especially near the solution, can significantly mitigate barren plateaus.The authors identify further exploration of these and other strategies as an important direction for future research.
Appendices
The appendices define the operator, Haar-measure, design, and frame-potential tools used to analyze cost gradients. They then derive unbiasedness and gradient-variance bounds, including extensions to generalized costs and diamond-norm expressibility.
- Definitions and identities: The appendices introduce Schatten norms, diamond-norm distance, Haar-measure properties, and symbolic integration identities for unitary ensembles.The diamond norm measures distinguishability between quantum operations, while the Haar measure supplies the uniform unitary reference.
- Designs and expressibility: An ansatz forms a t-design when its averaging agrees with Haar averaging through the t-th moment; the analysis specializes to t = 2.The second-moment superoperator is used to quantify expressibility relative to a 2-design.
- Frame potentials: State- and Hamiltonian-dependent frame potentials relate operator-dependent expressibility measures to quantities used in the gradient analysis.These expressions are used to evaluate the expressibility of different ansätze.
- Gradient averages: For the random layered ansatz, the average partial derivative of the generic cost vanishes, so the cost landscape is unbiased.The appendices rewrite the cost and derivative to expose dependence on the differentiated circuit rotation before averaging.
- Gradient-variance bounds: The variance of the partial derivative provides the starting point for deriving the main expressibility-dependent bounds.The derivation uses independent left and right unitary ensembles and then applies trace inequalities and Haar identities.
- Generalized costs: The gradient-variance results extend to generalized cost functions involving multiple states and Hamiltonian terms.The appendices derive corresponding expressions and bounds by substituting state- and Hamiltonian-dependent second-moment operators.
- Diamond-norm formulation: The bounds are also formulated using diamond-norm expressibility, a natural operational distance for ε-approximate t-designs.The appendix specifically derives bounds for k-local costs in this formulation.
G. Numerically studying the correlations between expressibility and cost partial derivatives
The numerical study varies ansatz expressibility through circuit depth, parameter correlations, and rotation restrictions, then compares gradient variance with frame-potential measures. It finds clear but imperfect correlations and shows that Hamiltonian-dependent measures capture cost locality.
- Experimental design: The study plots partial-derivative variance against ansatz expressibility while tuning circuit depth, parameter correlations, and rotation directions.These are the three principal ways used to vary expressibility numerically.
- Expressibility measures: Larger frame-potential ratios indicate more inexpressive ansätze, while ratios approaching 1 correspond to maximally expressive exact 2-designs.The plotted ratios compare the true frame potential with the Haar frame potential.
- Correlation results: The Spearman correlations between gradient variance and state- and Hamiltonian-dependent expressibility parameters are 0.78 and 0.80, with p-values of 1.19 × 10^-7 and 1.18 × 10^-7, respectively.The correlations combine the different methods used to tune expressibility.
- Locality dependence: The Hamiltonian-dependent frame potential captures locality effects that the state-dependent frame potential cannot capture.As depth increases, the Hamiltonian-dependent quantity decreases for local costs but remains effectively constant for global costs, matching their gradient-variance behavior.
- Correlation limits: The correlation between gradient variance and expressibility is clear but not perfect, consistent with the analytical quantities being upper bounds.The authors state that further work is needed to understand the detailed structure of this correlation.