Source-linked AI summary

Exploiting symmetry in variational quantum machine learning

Johannes Jakob Meyer, Marian Mularski, Elies Gil-Fuster, Antonio Anna Mele, Francesco Arzani, Alissa Wilms, Jens Eisert

arXiv:2205.06217v1quant-phcs.AIcs.LG

TL;DR

The paper addresses the limited guidance for choosing parametrizations in variational quantum learning by exploiting symmetries of the learning task. It constructs equivariant embeddings and gatesets using representation theory, reports better generalization on symmetric toy problems, and extends the approach to variational eigensolvers. The scope is bounded by representation and implementation pitfalls, including cases where symmetrization trivializes operations.

  • Problem

    Suitable parametrizations that encode task-relevant inductive bias remain poorly understood in variational quantum learning.

  • Method

    Representation theory transforms standard gatesets into equivariant gatesets and combines them with equivariant embeddings to construct invariant models.

  • Results

    Invariant models showed better generalization on tic-tac-toe and autonomous-vehicle toy problems, while equivariant ansätze often improved energy estimates, iteration counts, and barren-plateau behavior.

  • Takeaways & Limitations

    Symmetry-aware gatesets provide a blueprint for incorporating task structure into variational quantum models and other symmetric variational problems.

  • Takeaways & Limitations

    The approach has pitfalls, including generator trivialization, and may require using only subgroups rather than the full symmetry group.

Abstract

from arXiv · show

Variational quantum machine learning is an extensively studied application of near-term quantum computers. The success of variational quantum learning models crucially depends on finding a suitable parametrization of the model that encodes an inductive bias relevant to the learning task. However, precious little is known about guiding principles for the construction of suitable parametrizations. In this work, we holistically explore when and how symmetries of the learning problem can be exploited to construct quantum learning models with outcomes invariant under the symmetry of the learning task. Building on tools from representation theory, we show how a standard gateset can be transformed into an equivariant gateset that respects the symmetries of the problem at hand through a process of gate symmetrization. We benchmark the proposed methods on two toy problems that feature a non-trivial symmetry and observe a substantial increase in generalization performance. As our tools can also be applied in a straightforward way to other variational problems with symmetric structure, we show how equivariant gatesets can be used in variational quantum eigensolvers.

I. PRELIMINARIES

Variational quantum models use parametrized circuits, data embeddings, and measurements to produce predictions or optimize costs. The section introduces symmetry groups and shows how data embeddings can induce corresponding unitary representations for invariant quantum learning models.

  • Parametrized quantum circuits prepare states whose measured expectation values feed a cost function.
  • Data re-uploading interleaves a fixed data-embedding unitary with parametrized circuits to improve model expressivity.
  • A symmetry group acts on the Hilbert space through a unitary representation, while task predictions remain invariant under transformations of the data.
  • The framework targets discrete and continuous data symmetries and provides embeddings that support symmetry-invariant data re-uploading models.
  • Introductory example: The two-feature example has Z2 × Z2 symmetry, represented quantumly by qubit SWAP and simultaneous Pauli X operations.
  • An equivariant embedding satisfies a compatibility relation between data transformations and unitary conjugation on the Hilbert space.

Permutation symmetries

Permutation symmetries can be encoded with qubit permutations and Pauli rotations, with extensions to sign flips, higher-order features, and compact encodings. Greater compression generally increases implementation difficulty and can trivialize symmetric operations.

  • Permutation symmetries reshuffle data features, and every permutation can be built from transpositions.
  • The straightforward encoding assigns one qubit and one local RZ rotation to each feature, inducing permutations through qubit SWAP operations.
  • The same construction extends from permutations to sign flips, yielding the semidirect-product symmetry group Sd ⋊ Z2^d.
  • Equivariant embeddings trade qubit count against gate complexity, motivating intermediate constructions between one-local and highly controlled encodings.
  • Higher-order feature monomials can be embedded with multi-qubit Pauli-word generators while retaining equivariance under feature permutations.
  • Compact encodings can represent up to 2^n features, but their rotations are difficult to implement and may leave only trivial symmetry-respecting operations.

Continuous symmetries

The embedding framework also covers continuous symmetries by representing rotations and reflections on quantum systems. The construction realizes SO(3) and extends it to O(3) with an additional qubit.

  • Continuous data symmetries are described using Lie groups, which combine group structure with differentiability.
  • Representation theory guarantees that compact Lie groups can be represented within suitably large unitary groups.
  • SO(3): The paper constructs an equivariant embedding for SO(3) using a data-to-Bloch-sphere mapping and Pauli-rotation conjugation.
  • O(3): Adding a qubit realizes a fixed reflection alongside rotations, generating the full orthogonal group O(3).

The general case

The general construction embeds data into a quantum Lie algebra and matches the data representation with a restricted adjoint representation. This requires an appropriate finite-dimensional unitary representation, and classifying realizable embeddings remains open for more exotic symmetries.

  • General equivariant embeddings require the symmetry group to have non-trivial finite-dimensional unitary representations.
  • The data embedding is modeled as a linear map from data space into su(2^n), with symmetries represented on the data by a Lie-group action.
  • The induced Lie-algebra transformation is implemented by unitary conjugation, namely the adjoint action of the symmetry representation.
  • An equivariant embedding exists when a subspace of su(2^n) carries the same representation as the symmetry Lie algebra acting on the data.
  • A matching Lie subalgebra alone is insufficient; the required subrepresentation must occur inside the adjoint representation of su(2^n).

III. GATE SYMMETRIZATION

Gate symmetrization uses representation theory to transform standard ansatz generators into operators commuting with the problem’s symmetry representation. This produces equivariant gatesets while potentially reducing ansatz complexity and expressivity.

  • Motivation: Symmetries can guide variational model parametrizations by encoding task-relevant data and reducing free parameters.The same symmetry-based construction is also relevant to variational ground-state searches, where expressivity in relevant Hilbert-space regions is the goal.
  • Definition: An equivariant gate is defined by commuting with every operator in the symmetry representation.For an exponential gate generated by G, equivariance holds if and only if G commutes with all representation operators.
  • Trade-offs: Equivariance is advantageous but not universal because full-group symmetrization can reduce the available operations and the ansatz’s expressivity.The authors therefore note that subgroups may sometimes be preferable.
  • Construction: Twirling projects a generator onto the operators that commute with the symmetry representation.For finite groups it uses a uniform average; for Lie groups it uses integration over the Haar measure.
  • Construction: Applying twirling to the generators of a standard gateset yields an equivariant gateset that can be used to construct symmetry-respecting trainable blocks.The resulting gateset can also include non-parametrized gates through direct symmetrization or fixed-angle parametrized forms.

Ansatz symmetrization

Symmetrizing an ansatz replaces its generators with symmetry-compatible counterparts, often collapsing equivalent local operations. The resulting gateset respects the selected symmetries but may have substantially reduced expressivity.

  • Ansatz conversion: A complete ansatz can be converted into an equivariant ansatz by replacing its gateset with a precomputed equivariant counterpart.For practical purposes, the equivariant gateset can usually be computed efficiently beforehand.
  • Combined symmetries: For the two-qubit example, the full symmetry group is Z2 × Z2, generated by SWAP and the sign flip X1X2.The construction first symmetrizes over the two subgroups and then combines the results.
  • Exchange symmetry: Under exchange symmetry, local generators on corresponding qubits map to shared symmetrized operators, reducing gateset cardinality and parameter count.For example, X1 and X2 symmetrize to the same operator, while Z1Z2 already commutes with SWAP.
  • Sign-flip symmetry: Under sign-flip symmetry, only local Pauli X rotations remain among the considered local gates.The other local Pauli generators change sign under conjugation by X1X2 and are removed by symmetrization.
  • Trade-offs: Imposing the full symmetry greatly reduces available operations, with a corresponding cost in ansatz expressivity.This reduction motivates treating symmetry selection as an expressivity trade-off rather than an unconditional benefit.

Pitfalls to avoid

The construction of invariant quantum learning models combines equivariant embeddings, equivariant trainable blocks, an invariant initial state, and an invariant observable. Its main limitation is the trade-off between symmetry specialization, expressivity, universality, and hardware cost.

  • Expressivity trade-off: Symmetrization always trades increased equivariance and specialization for reduced ansatz expressivity.The paper recommends considering subgroups or limited symmetry-breaking gates when full symmetry is too restrictive.
  • Generator trivialization: Symmetrization can trivialize generators, leaving no gates even when the original parametrization was universal.For the representation U0 = I and U1 = X, both Y and Z twirl to zero, although RX rotations remain symmetry-compatible.
  • Universality: Depending on the starting gates, symmetrization can destroy universality because the resulting gates may not generate all symmetry-compatible unitaries.The paper notes that quantum-chemistry mappings require 3-local unitaries for all symmetry-preserving transformations.
  • Embedding requirement: The data embedding must be equivariant so the learning-task symmetry is meaningfully represented in Hilbert space.Without an equivariant embedding, the authors state that constructing invariant quantum learning models is unlikely.
  • Invariant models: Invariant predictions arise by combining an equivariant circuit with an invariant initial state and an invariant observable.The model is built from alternating data embeddings and trainable blocks before evaluating the final observable.

V. NUMERICAL EXPERIMENTS

Numerical experiments compare invariant and non-invariant models on toy learning tasks and evaluate equivariant ansätze on ground-state problems. Invariant models generalize better despite lower training performance, while ground-state benefits are more mixed.

  • Learning experiments: Invariant learning models under-perform on training data but achieve much better generalization performance than non-invariant models.The experiments compare models on two selected toy problems with non-trivial symmetry.
  • Ground-state experiments: Equivariant ansätze produce better ground-state approximations on average and require fewer iterations to converge across three studied spin models.The models are the transverse-field Ising, Heisenberg, and longitudinal-transverse-field Ising models.
  • Ground-state experiments: Equivariant ansätze can mitigate barren plateaus, but their applicability to ground-state problems is less clear and can have downsides.The paper discusses these downsides in connection with the broader expressivity trade-off.
  • Tic-tac-toe: Tic-tac-toe classification tests whether models distinguish cross wins, circle wins, and draws while respecting the board’s rotational and reflection symmetries.The task’s symmetry group is the dihedral group D4 of order 8.
  • Tic-tac-toe: The tic-tac-toe experiment primarily demonstrates an end-to-end implementation of symmetrization and compares equivariant with standard gatesets.The authors note that the task itself is easily solved by a classical deterministic algorithm.

Dataset

The paper constructs symmetry-aware datasets and variational models for tic-tac-toe and autonomous-driving scenarios, using equivariant embeddings and shared gates. Across experiments, invariant models trade training-set expressivity for consistently stronger test performance.

  • Tic-tac-toe: Tic-tac-toe games are encoded on nine qubits, with board rotations and reflections forming the symmetry group D4.Each board field maps to a vector value, while labels represent cross wins, circle wins, or draws.
  • Tic-tac-toe: Pauli-X rotations encode the three field values using multiples of 2π/3, and symmetry acts by permuting equivalent qubits.Corners, edges, and the middle form equivalence classes under the board symmetries.
  • Tic-tac-toe: The tic-tac-toe circuit initializes |0⟩, repeatedly re-uploads data, applies shared parametrized layers, and predicts from three invariant observables.The default layer combines single-qubit gates with entangling gates, and the largest measured expectation selects the class.
  • Results: Invariant models match or underperform non-invariant models on training accuracy but consistently achieve higher test accuracy.The comparison uses models without symmetry-imposed parameter sharing as the more expressive baseline.
  • Results: Across circuit sizes and randomized trainable layouts, reduced expressivity produces similar performance on seen and unseen data, while non-invariant models show overfitting.The size sweep excludes five high-resource architecture pairs because of limited computational resources.

Learning model

The learning models encode rotationally symmetric driving scenarios with Z4-equivariant data re-uploading circuits and compare invariant models against more expressive non-invariant variants. In both autonomous-driving experiments, invariant models generally sacrifice training accuracy while improving test generalization.

  • Model construction: The autonomous-driving toy model has symmetry group Z4 because the simplified scenario representation retains rotations but lacks mirror symmetry.Scenes are encoded through Pauli-X rotations whose angles are multiples of 2π/3 of the array values.
  • Model construction: Predicted difficulty is obtained from a normalized Pauli-Z expectation value and rounded to the nearest level in {0, 0.2, 0.4, 0.6, 0.8, 1}.The training objective is an l2-loss over encoded scenarios and difficulty labels.
  • Evaluation: The autonomous-driving experiments compare invariant circuits with non-invariant versions lacking parameter sharing, which have higher expressivity.Performance is evaluated using training and test classification accuracy.
  • Results: Across most architecture sizes, invariant models perform worse on training data but better on test data than non-invariant models.The results support a bias-variance interpretation in which non-invariant models overfit while invariant models occupy a better-performing regime.
  • Results: With randomized trainable-block orderings, invariant models outperform non-invariant models in almost every tested circuit construction.Each randomized layer is repeated three times, and the experiment varies the spelling of the trainable circuit components.
  • Transverse-field Ising model: For the TFIM, the equivariant QAOA ansatz outperforms the non-equivariant ansatz in both energy and convergence iterations when p ≥ N/2.For p < N/2, the equivariant ansatz is outperformed because its restricted expressivity cannot reach the ground state.

Heisenberg model

The Heisenberg example applies symmetry-preserving gate construction to a continuous SU(2)-symmetric ground-state problem. Symmetrization reduces parameters and yields faster optimization at every depth, while energy performance favors the non-equivariant ansatz at small depths and the equivariant ansatz at sufficiently large depths.

  • Symmetry: The Heisenberg Hamiltonian is invariant under joint rotations of all spins, giving it an SU(2) symmetry.The model describes the alignment of neighboring spins, which is unchanged by simultaneous global rotations.
  • Ansatz construction: The variational state is initialized in the zero-total-spin symmetry sector before applying the Heisenberg ansatz.This initialization places the state in the symmetry sector relevant to the ground state.
  • Ansatz construction: The starting non-equivariant ansatz has seven parameters per layer: three for each anisotropic even and odd Hamiltonian, plus one Pauli-Y rotation parameter.The corresponding unitaries can be decomposed into two-qubit gates because the Hamiltonian terms commute within each sublattice.
  • Symmetrization: Gate symmetrization uses a 2-design twirl to remove non-symmetric components and produce the equivariant ansatz.The twirl reduces relevant generators to symmetry-compatible combinations, including particle-number-conserving Givens rotations.
  • Symmetrization: The equivariant Heisenberg ansatz has 2p parameters, compared with 7p for the non-symmetric ansatz.The construction uses isotropic even and odd Hamiltonians with β = γ = (1, 1, 1).
  • Results: The equivariant ansatz reaches minimum energy faster across all depths but outperforms the non-equivariant ansatz in energy only at sufficiently large depth.At small depth, the non-equivariant ansatz benefits from greater expressivity; once the equivariant ansatz is sufficiently expressive, the ordering reverses.
  • Optimization diagnostics: The barren-plateau comparison evaluates derivative variance for TFIM and Heisenberg circuits using specified depths, observable, and 1000 random parameter samples.The derivative is taken with respect to the first-qubit, first-layer rotation angle using Z1Z2.

Barren plateaus

The paper examines whether symmetry-aware ansätze affect barren plateaus and variational eigensolver performance. In the LTFIM lattice experiment, equivariant circuits outperform non-equivariant circuits at tested depths, but the advantage weakens at larger depth.

  • Barren plateaus: Symmetrization reduces free parameters and changes the dynamical Lie algebra, both of which are relevant to barren plateaus.The authors therefore analyze barren-plateau behavior in transverse-field Ising and Heisenberg model setups.
  • LTFIM lattice: The LTFIM lattice has geometric symmetries matching the tic-tac-toe task, with three independent families of qubits and edges.Equivalent sites and edges are grouped according to the lattice symmetry.
  • Results: For smaller circuit depths, the equivariant LTFIM ansatz achieves lower mean energy and requires fewer iterations than the non-equivariant ansatz.Figure 13 averages ten random initializations for each layer count p.
  • LTFIM ansatz: Equivariant circuits share gate parameters across symmetry-equivalent edges and spins, reducing each p-layer circuit from (16 + 9 + 9)p to (3 + 3 + 3)p free parameters.The invariant state vector |+⟩⊗N is used for initialization.
  • Results: The equivariant advantage is not stable at larger depths, where the achieved energy increases for the equivariant ansatz.The authors caution that symmetrization is helpful but not a universal solution.

Discussion

The discussion attributes faster optimization in some cases to reduced parameter counts, while noting a trade-off between symmetry enforcement and expressivity. The paper concludes that symmetry-aware embeddings and equivariant gatesets provide a construction blueprint, but broader theory and larger-scale tests remain open.

  • Discussion: Equivariant ansätze can require fewer optimizer iterations because reducing free parameters simplifies the optimization problem.The authors report this improvement across some of their numerical experiments.
  • Discussion: Minimal achievable energy shows a trade-off between the gain in equivariance and the reduction in expressivity.At small depths, non-equivariant ansätze may explore more of Hilbert space; at larger depths, restricting to the relevant subspace can help.
  • Discussion: The authors motivate carefully engineered symmetry-breaking alongside equivariant ansätze, especially at small circuit depths.They cite shallow parity-preserving ansätze as potentially obstructing ground-state preparation for certain Hamiltonians.
  • Summary and outlook: Representation-theoretic equivariant gatesets enable invariant re-uploading models built from equivariant embeddings, equivariant trainable blocks, invariant initialization, and invariant observables.This provides a blueprint and tools for constructing invariant variational quantum machine-learning models.
  • Results: Experiments on tic-tac-toe and autonomous-vehicle tasks found better generalization for invariant models, including comparisons with randomly constructed invariant architectures.The reported comparisons were designed to avoid cherry-picking.
  • Beyond machine learning: The same gateset construction can be applied to variational eigensolver problems with symmetries arising from conserved Hamiltonian quantities, although results there are less clear.The paper demonstrates this direction on transverse-field Ising, Heisenberg, and related models.
  • Limitations and outlook: The study examines equivariant quantum embeddings for O(3), while other Lie-group symmetries remain a future research direction.The authors also call for rigorous generalization results and note that learning experiments used at most nine qubits because of computational limits.

Appendix A: Additional figures

The appendix figures sweep circuit depth and layer count for two grid-structured classification tasks. They plot mean test accuracies for invariant and non-invariant models, with darker colors showing averages over ten runs.

  • Tic-tac-toe: Figure 14 sweeps l and p for tic-tac-toe classification and plots mean test accuracy for invariant and non-invariant models.Darker colors represent the average over ten runs; each model uses l layers containing data encoding and p repetitions of “cemoid”.
  • Autonomous vehicles: Figure 15 performs the corresponding l-and-p sweep for autonomous-vehicle scenarios arranged in a grid.It likewise compares mean test accuracies for invariant and non-invariant models, with darker colors indicating ten-run averages.
Loading 2205.06217v1…