Source-linked AI summary

Trainability of Dissipative Perceptron-Based Quantum Neural Networks

Kunal Sharma, M. Cerezo, Lukasz Cincio, Patrick J. Coles

arXiv:2005.12458v2quant-phcs.LG

TL;DR

The paper analyzes gradient behavior in dissipative quantum neural networks for quantum machine-learning tasks. It proves zero average cost-function derivatives and derives variance upper bounds under specified circuit and perceptron assumptions.

  • Problem

    The paper addresses how cost-function gradients behave in dissipative quantum neural networks applied to quantum machine-learning tasks.

  • Method

    The analysis proves gradient properties for random parameterized circuits and parameter-matrix-multiplication updates, using 2-design assumptions and extending the setting to perceptrons acting on n + m qubits.

  • Results

    The average partial derivative of the cost function is zero, and its variance admits upper bounds for circuit parameters and time-step parameters under the stated DQNN constructions.

  • Takeaways & Limitations

    The formalism and results can be extended to other supervised quantum-machine-learning tasks and to DQNNs with multiple output qubits.

  • Takeaways & Limitations

    The proofs initially assume each perceptron acts on one output qubit and that the DQNN has no hidden layers, with generalizations treated separately.

Abstract

from arXiv · show

Several architectures have been proposed for quantum neural networks (QNNs), with the goal of efficiently performing machine learning tasks on quantum data. Rigorous scaling results are urgently needed for specific QNN constructions to understand which, if any, will be trainable at a large scale. Here, we analyze the gradient scaling (and hence the trainability) for a recently proposed architecture that we called dissipative QNNs (DQNNs), where the input qubits of each layer are discarded at the layer's output. We find that DQNNs can exhibit barren plateaus, i.e., gradients that vanish exponentially in the number of qubits. Moreover, we provide quantitative bounds on the scaling of the gradient for DQNNs under different conditions, such as different cost functions and circuit depths, and show that trainability is not always guaranteed.

Supplemental Material for Trainability of Dissipative Perceptron-Based

The supplemental material establishes the proof framework for the main results, using Haar-measure properties, unitary-design definitions, and trace identities. It also states simplifying assumptions and extensions for the DQNN analysis.

  • Proof framework: The proofs establish that the average cost-function gradient is zero for the considered DQNNs.The supplemental material identifies this as a preliminary result used in proving the main theorems.
  • Assumptions and extensions: The analysis initially assumes one output qubit per perceptron and no hidden layers, then generalizes to multiple output qubits and hidden layers.These extensions are treated in later supplemental sections.
  • Mathematical tools: The supplemental material defines unitary t-designs through agreement between finite-unitary averages and Haar averages for degree-t polynomials.The definition applies to polynomials in unitary matrix elements and their conjugates.
  • Mathematical tools: Haar integration supplies the first- and second-moment identities used to evaluate traces involving random perceptrons.The identities are applied to linear operators and bipartite Hilbert spaces, with partial traces handling subsystem structure.
  • Mathematical tools: The proofs use partial-trace lemmas for multipartite operators to reduce Haar-averaged expressions involving perceptron layers.These lemmas define reduced operators through traces over selected subsystems and support the theorem derivations.
  • Assumptions and extensions: For simplicity, the rigorous proofs use computational-basis tensor-product output states and DQNN unitaries acting on n + 1 qubits, with broader cases argued separately.The supplemental material states that arbitrary tensor-product states and additional layers are treated through generalization arguments.

B. Generalization of our results to quantum machine learning task

The supplemental section extends the DQNN formalism from state-preparation tasks to supervised quantum machine-learning tasks with labeled quantum inputs. It describes binary-label prediction through measurements on the output state.

  • Problem formulation: The supervised task uses training pairs {|ψin_x⟩, yx}, where the DQNN predicts a label matching the classical label yx.The section specializes to binary labels yx ∈ {−1, 1}.
  • Prediction procedure: Predicted labels are obtained by performing a binary measurement on the DQNN output states.The measurement is taken in the z basis, producing bitstring outcomes on the measured output qubits.
  • Generalization: The state-preparation formalism extends directly to supervised quantum machine-learning tasks.The supplemental construction states that the supervised-learning training-set form matches the main-text formalism.

C. Proof of ⟨∂C⟩= 0

This section proves that the average partial derivative of the DQNN cost is zero, considering both random parameterized circuits and parameter matrix multiplication. The result follows from averaging over random perceptrons and trace identities.

  • Main result: ⟨∂C⟩ = 0 for the considered DQNN cost function, so the average gradient is not biased toward any particular value.The proof covers both random parameterized quantum circuits and parameter matrix multiplication.
  • Random parameterized circuits: For random parameterized circuits, independent random perceptrons and a one-design condition yield a vanishing averaged partial derivative.The derivation uses independence assumptions and trace expressions involving the perceptron acting on the differentiated parameter.
  • Cost functions: The cost function is defined for input-output quantum-state training pairs using either a global observable or a local output-layer operator.The local operator acts nontrivially on one output qubit while identities act on the remaining output qubits.
  • Parameter matrix multiplication: The parameter-matrix-multiplication method updates perceptrons through a time parameter and uses Hermitian generators in the derivative expression.The perceptrons are randomly initialized and updated at each step s.
  • Parameter matrix multiplication: At the initial time step, independent random initialization leads to the same unbiased-gradient conclusion for the parameter-matrix-multiplication setting.The proof evaluates the averaged derivative using the propagated operators and the initial perceptron ensemble.

D. Proof of Theorem 2

The proof of Theorem 2 reduces gradient-variance analysis to second moments because the mean gradient is zero. It treats global and local cost functions under simplifying architectural assumptions.

  • Theorem statement: Theorem 2 bounds the variance of a partial derivative with respect to the time-step parameter for deep global perceptrons initialized as independent 2-designs.The theorem covers both global and local cost operators.
  • Proof assumptions: The proof assumes no hidden layers and n qubits in both input and output layers before considering broader cases elsewhere.Randomly initialized perceptrons are denoted without their layer superscripts in this derivation.
  • Variance calculation: Because ⟨∂C⟩ = 0, the gradient variance depends only on the second moment of the partial derivative.The proof analyzes squared trace terms and their cross terms over training examples and perceptrons.
  • Variance calculation: The proof bounds individual squared trace contributions and then bounds cross terms to obtain the theorem’s variance scaling.The relevant terms are indexed by training examples x and perceptrons j.

1. Global Cost

This section derives an upper bound on gradient-related terms for a global cost function by recursively averaging over randomly initialized perceptrons modeled as 2-designs. The resulting bound is obtained through operator decompositions, trace inequalities, and recursive relations.

  • Global-cost variance: The global-cost analysis estimates the variance of a single gradient term with fixed x and j.The derivation begins by invoking Lemma 1 and analyzing the squared term associated with the global operator.
  • Operator construction: The operators A(x,j) and B(x,j) are expressed through partial traces and matrix elements of the input state and perceptron transformations.The construction uses traces over output qubits and bitstring-indexed matrix elements.
  • Recursive averaging: An upper bound on the average of (s(x,j)r(x,j))^2 is established when Vj+1 forms a 2-design.The bound is then used as the starting point for recursive averaging across later perceptrons.
  • Recursive averaging: Because the averaged state remains a quantum state, the same construction can be recursively applied over Vj+2 through Vn.The recursion uses the definitions of the subsequent s and r operators and the assumption that all randomly initialized perceptrons form 2-designs.
  • Combining bounds: The lower-layer contribution is likewise averaged over V1 through Vj−1, producing the final bound after combining the derived relations.The calculation proceeds through the corresponding A(x,j) and r(x,j) terms before using the normalization relation in the final step.

b. Fixed j and different x

This subsection bounds cross terms with equal j and different input labels x. The derivation combines recursive averaging with Cauchy–Schwarz-based estimates to obtain the resulting bound.

  • Cross-term structure: The cross terms considered have equal j but different x.The subsection explicitly targets terms of the form involving distinct x labels at the same layer index.
  • Cross-term bound: Cauchy–Schwarz is invoked together with earlier relations to derive an upper bound on the cross terms.The bound is assembled after estimating the relevant intermediate expressions.
  • Recursive averaging: The calculation averages s(x,j) and recursively relates r(x′,j) to later-layer quantities.These relations are evaluated with respect to the next perceptron and then propagated recursively.
  • Cross-term bound: Combining the preceding bounds yields the subsection’s final estimate for these equal-j, different-x contributions.The conclusion follows by combining the two displayed bounds established in the subsection.

c. Different j and different x

This subsection analyzes cross terms involving different layer indices and different input labels. It separates cases by index ordering and uses Lemma 2, Cauchy–Schwarz, and local-cost operator structure to bound or eliminate contributions.

  • Cross-term structure: The proof considers cross terms with different j and different x, assuming without loss of generality that j < k.The terms are organized according to the relative positions of the layer indices and the associated perceptron products.
  • Lemma-based reduction: Lemma 2 is applied after identifying the relevant perceptron products with the operators appearing in the lemma.The mappings include U, S, S′, H, K, P, and P′.
  • Combined result: Combining the results from the separate cases produces the subsection’s overall bound for different-j, different-x terms.The final estimate is stated after the case-by-case analysis.
  • Local-cost cases: For local cost functions, the variance analysis is reduced to cases in a triple summation involving the relative ordering of i and j.The cases i < j, i = j, and i > j are treated separately.
  • Local-cost cases: When i < j, the relevant trace term vanishes after combining the structural relations with the preceding bounds.The derivation uses the relations associated with the support of the operators and the preceding inequalities.
  • Local-cost cases: The i = j and i > j cases are controlled through positivity, operator-spectrum bounds, and Cauchy–Schwarz estimates.These steps bound the corresponding contributions before they are combined with the result for the remaining index configurations.

c. Different j, different i, and different x

This subsection combines the preceding cross-term analyses for different layer indices, different operator indices, and different input labels. The resulting estimate uses the same function f(n) introduced earlier.

  • Final cross terms: The remaining cross terms are treated under the assumption j < k and with σout allowed to take any form covered by the preceding proof.The argument transfers the result from the earlier subsection to the present operator configuration.
  • Combined result: Combining the results from Sections D 2 a–D 2 c yields the final bound for this class of terms.The result is stated after the preceding cases have been assembled.
  • Combined result: The bound is expressed using the same f(n) defined in (D1).The subsection therefore reuses the earlier n-dependent function rather than introducing a new scaling function.

E. Proof of Theorem 1

The proof bounds gradient-variance contributions for global and local cost functions by averaging over randomly initialized perceptrons and separating diagonal and cross terms. It treats single-output contributions, cross terms, and cases determined by output-qubit ordering.

  • Theorem assumptions: Theorem 1 assumes deep global perceptrons whose unitaries form independent 2-designs over n + 1 qubits.The theorem concerns gradient variance for parameters in a perceptron under global or local cost operators.
  • Proof strategy: The proof first analyzes global-cost gradient variance and then derives bounds for local-cost gradient variance.The derivation separates the two cost-function settings before combining the resulting bounds.
  • Global costs: For global costs, the proof bounds single-x terms by recursively averaging over randomly initialized perceptrons and evaluating associated operator expressions.The construction uses operators on all input qubits plus a selected output qubit and traces over the remaining output subsystem.
  • Global costs: Cross terms with different x values are bounded separately using operator identities and recursive applications of the same averaging arguments.The proof combines bounds for these cross terms with the single-term estimates.
  • Local costs: For local costs, the proof divides contributions according to whether the measured output index i is less than, greater than, or equal to the differentiated perceptron index j.The i > j case differs because the perceptron unitaries do not commute.
  • Local costs: The local-cost analysis recursively applies bounds across the three index cases and then combines diagonal and cross-term estimates.The resulting bounds are assembled from the case-specific inequalities and recursive lemmas.

d. Different x and different i.

This subsection bounds local-cost cross terms involving different input labels and output indices by partitioning the cases according to their relation to the differentiated perceptron index. The case analysis is then combined into a final bound.

  • Case analysis: Local-cost cross terms are organized by whether either output index is smaller than j, whether i = j and i′ > j, or whether both exceed j.Each ordering produces a distinct operator and bitstring analysis.
  • Case analysis: When i = j and i′ > j, the proof shows that the relevant operator difference is independent of the bitstrings p and p′.This independence yields a simplified bound for that cross-term case.
  • Combined bound: The bounds for the separate index-order cases are combined to obtain the subsection’s final local-cost cross-term bound.The combination uses recursive applications of the preceding lemmas.

F. DQNNs with unitaries acting on n + m qubits

The analysis generalizes DQNN gradient-variance results from perceptrons acting on n + 1 qubits to unitaries acting on n input and m output qubits. It then extends the averaging argument across output blocks.

  • Generalized architecture: The generalized architecture uses perceptrons acting on n qubits from layer l and m qubits from layer l + 1, with Fig. 4 illustrating n = m = 2.Each perceptron can be represented by a unitary acting on n + m qubits.
  • Proof construction: The proof groups the m output qubits into blocks and defines reduced states and operators on the input layer plus one selected output block.Partial traces remove the remaining output qubits from the relevant expressions.
  • Assumptions: The generalized bounds assume the relevant operators satisfy a Hilbert-Schmidt norm condition, including Tr[(Hj)^2] ≤ 2^(n+m).This assumption enters the expectation-value bound for the generalized operator.
  • Results: The section states that the resulting analysis applies to both global cost functions and can be generalized similarly to local cost functions.The global result is obtained by combining the generalized intermediate bounds.

G. DQNNs with hidden layers

The hidden-layer analysis extends the preceding gradient-variance derivations to DQNNs with L hidden layers. It partitions qubits into index sets to remove unaffected degrees of freedom and recover analogous bounds.

  • Generalization: The section generalizes the results from earlier sections to DQNNs with L hidden layers.The same gradient-variance framework is applied to parameters in an interior-layer perceptron.
  • Results: The resulting global-cost derivation follows the earlier single-layer proof and yields an analogous bound for hidden-layer DQNNs.The text also states that the local-cost proof can be generalized in the same manner.
  • Index-set construction: For a parameter in V_j^l, the proof partitions qubits into four sets based on layer location and their indices relative to j.The sets distinguish earlier systems, the preceding hidden layer, earlier qubits in layer l, and later qubits.
  • State reduction: The action of preceding unitaries allows the state around the differentiated perceptron to be expressed using these index sets.The resulting expressions impose matching, Pauli-label, or zero constraints on the corresponding bitstring indices.
  • State reduction: Because the cost operator acts trivially on set S1, summing over its matching indices removes those qubit indices from the calculation.The sum produces the identity on S1, reducing the effective operator structure.
Loading 2005.12458v2…