Source-linked AI summary
Group-Invariant Quantum Machine Learning
Martin Larocca, Frederic Sauvage, Faris M. Sbahi, Guillaume Verdon, Patrick J. Coles, M. Cerezo
TL;DR
QML models with weak inductive biases face trainability and generalization concerns, motivating models that encode dataset symmetries. The paper develops a theoretical framework for constructing group-invariant QML models and applies it to continuous and discrete symmetry groups. Across these tasks, the framework recovers known quantum algorithms and identifies invariant models for several classification problems.
Problem
QML models with little or no inductive bias can have trainability and generalization issues, motivating schemes that encode available problem structure.
Method
The paper derives propositions characterizing QML models invariant under a dataset’s symmetry group and applies them to continuous Lie groups and discrete symmetry groups.
Results
The framework recovers known protocols and algorithms and yields invariant models for purity, time-reversal, multipartite entanglement, and graph-isomorphism classification tasks.
Takeaways & Limitations
The results provide groundwork for a more geometric and group-theoretic approach to QML model design.
Abstract
from arXiv · showhide
Quantum Machine Learning (QML) models are aimed at learning from data encoded in quantum states. Recently, it has been shown that models with little to no inductive biases (i.e., with no assumptions about the problem embedded in the model) are likely to have trainability and generalization issues, especially for large problem sizes. As such, it is fundamental to develop schemes that encode as much information as available about the problem at hand. In this work we present a simple, yet powerful, framework where the underlying invariances in the data are used to build QML models that, by construction, respect those symmetries. These so-called group-invariant models produce outputs that remain invariant under the action of any element of the symmetry group $\mathfrak{G}$ associated to the dataset. We present theoretical results underpinning the design of $\mathfrak{G}$-invariant models, and exemplify their application through several paradigmatic QML classification tasks including cases when $\mathfrak{G}$ is a continuous Lie group and also when it is a discrete symmetry group. Notably, our framework allows us to recover, in an elegant way, several well known algorithms for the literature, as well as to discover new ones. Taken together, we expect that our results will help pave the way towards a more geometric and group-theoretic approach to QML model design.
I. INTRODUCTION
The paper imports symmetry-based inductive biases from geometric deep learning into QML, defining dataset symmetries and models that respect them by construction. It develops a framework for symmetry-informed QML design and applies it across quantum classification settings.
- Motivation: Geometric deep learning views successful architectures as models whose inductive biases respect domain structure and symmetries.Such biases restrict the function space and can improve generalization, data efficiency, and optimization landscapes.
- Motivation: QML models with uninformed inductive biases can suffer trainability and generalization problems, while sharper priors narrow the effective search space.The paper identifies barren plateaus as one issue associated with highly expressive, problem-agnostic models.
- Framework: The paper characterizes QML models invariant under a dataset symmetry group G through Propositions 1–4 and applies the framework to several classification tasks.The applications include purity, time-reversal dynamics, multipartite entanglement, and graph isomorphism.
- Framework: A dataset symmetry group G consists of unitary operations that leave the labels unchanged, motivating models satisfying h(VρV†) = h(ρ).The paper proposes enforcing this invariance by construction rather than learning it solely through data augmentation.
- Framework: The framework distinguishes tasks with one shared symmetry group from binary tasks whose classes have separate groups G0 and G1.For separate class symmetries, models may be invariant under G0, G1, or both.
- Experimental settings: The paper analyzes conventional and quantum-enhanced experiments, where the latter permits coherent operations on multiple copies of each quantum state.The model structure is constrained by how quantum data can be stored, accessed, and measured.
E. Useful definitions
The framework uses commutants, higher-order symmetries, Lie algebras, and orthogonal complements to derive conditions guaranteeing group-invariant QML models. These conditions extend to multiple symmetry groups and distinguish invariance under each group or under their combination.
- Group symmetries: The commutant C(G) contains matrices commuting with every group element, while C^(k)(G) generalizes this condition to k-fold tensor representations.Hermitian representatives can be associated with higher-order symmetries, allowing them to serve as observables.
- Lie-group structure: For Lie groups, an associated Lie algebra g provides additional structure for constructing invariant models.The paper also defines the Hilbert–Schmidt orthogonal complement g⊥, which is not itself a Lie algebra.
- Invariant-model conditions: A Hypothesis Class 1 model is G-invariant when its effective observable eO(θ) belongs to the k-th-order commutant C^(k)(G).This follows because membership in the commutant makes the effective observable commute with V⊗k for every V in G.
- Invariant-model conditions: For Lie groups, invariance can alternatively be guaranteed by choosing the effective observable orthogonal to transformed k-copy states.The construction uses conditions involving the Lie algebra and its orthogonal complement, with assumptions on the identity operator and state support.
- Multiple symmetry groups: With two class-specific symmetry groups G0 and G1, commutant conditions determine whether a model is invariant under G0, G1, or both.The same logic extends to Lie groups through their respective Lie algebras and orthogonal complements.
- Multiple symmetry groups: The propositions extend beyond two classes to more general collections of symmetry groups, including multiclass classification.The model can be required to satisfy invariance conditions for all groups or selected groups.
IV. LIE GROUP-INVARIANT MODELS
The section applies the framework to Lie-group symmetries, beginning with purity classification under U(d). It proves that conventional single-copy invariant models cannot classify purity, motivating multi-copy quantum-enhanced models.
- Applications: The section applies the general invariant-model results to purity, time-reversal, and multipartite-entanglement datasets with Lie-group symmetries.The results are presented as theorems, with most proofs included constructively.
- Purity dataset: For the purity dataset, U(d) is the symmetry group because unitary transformations preserve quantum-state spectra and therefore purity.
- Conventional experiments: No conventional single-copy model in H1 can be both U(d)-invariant and informative for purity classification.The proof excludes the invariant models generated by the relevant propositions and then shows that no other invariant H1 models exist.
- Conventional experiments: U(d)-invariant models from the commutant reduce to constant predictions, so they provide no information about purity.Schur’s lemma yields eO(θ) = λ11, giving h(1)_θ(ρ) = λ for every state.
- Conventional experiments: Purity classification requires nonlinear dependence on ρ, whereas the considered single-copy models are linear and cannot linearly separate pure from mixed states.The section explains that purity is a second-order polynomial in the matrix elements of ρ.
2. Quantum-enhanced experiments
Quantum-enhanced experiments use multiple coherent copies of a state, allowing invariant observables built from identity and permutation operators. Two copies suffice for perfect purity classification, while distinct class symmetries enable a time-reversal classification strategy with possible noise limitations.
- Quantum-enhanced experiments: Two coherent copies are sufficient to classify states by purity in a quantum-enhanced experiment.
- Purity dataset: A U(d)-invariant two-copy model with a nonzero SWAP component can perfectly classify the purity dataset.The general construction guarantees suitable networks and observables, and any choice with a2 ≠ 0 yields perfect classification.
- Quadratic symmetries: For U(d), quadratic invariant observables are spanned by the two-copy identity and SWAP, reflecting the two elements of the k = 2 symmetric group.The symmetric group acts by permuting subsystems of tensor-product copies.
- Implementation: Constructing the required invariant observable does not by itself ensure efficient measurement: directly measuring SWAP can require exponentially many Pauli terms.An ancilla and the Hadamard test are given as an efficient route for estimating SWAP.
- Time-reversal dataset: In the time-reversal dataset, label-0 states have U(d) symmetry while label-1 states have O(d) symmetry, enabling models invariant under one class symmetry but not the other.O(d) preserves the time-reversal symmetry of the label-1 states.
- Time-reversal dataset: The time-reversal classification criterion is perfect when c lies outside [b1, b2], but may be noisy when c lies inside that interval.The latter case can produce misclassification because different classes may yield the same prediction.
1. Conventional experiments
In conventional experiments, first-order invariant models cannot generally distinguish time-reversal states, while suitable constructions can achieve noisy classification. However, their predictions may concentrate near zero, requiring exponentially many repetitions.
- No non-trivial linear symmetries of O(d) are available for classifying the time-reversal dataset in conventional experiments.
- Real-valued quantum neural networks can produce O(d)-invariant models that perform noisy classification of time-reversal data.
- Quantum-enhanced experiments: Theorem 5 shows that time-reversal-symmetric dynamics can be perfectly classified using O(d)-invariant models with a suitable two-copy initial state.The special choice |Ψin⟩ = |Φ+⟩ recovers an existing algorithm.
- Quantum-enhanced experiments: O(1) experiment repetitions suffice for dynamics classification, contrasting with the exponential repetition requirement for time-reversal-state classification.
C. Multipartite entanglement dataset
The multipartite entanglement dataset classifies pure quantum states by whether their entanglement exceeds a positive threshold. Local-unitary invariance supplies the shared symmetry group for both classes.
- Multipartite entanglement is difficult to characterize because its complexity scales exponentially with the number of parties and lacks a unique measure.
- The dataset separates states with multipartite entanglement measure E(ρ) = b > 0 from separable states with E(ρ) = 0.
- The shared symmetry group is G = N_n j=1 su(2), because local unitaries leave multipartite entanglement unchanged.
1. Conventional experiments
For multipartite entanglement, conventional one-copy invariant models are trivial and cannot classify the data, whereas two-copy models yield non-trivial entanglement measures and perfect classification.
- No conventional one-copy invariant model can provide relevant information for classifying the multipartite entanglement dataset.
- Conventional experiments: One-copy invariant constructions yield constant predictions, such as h(1)_θ(ρ_i) = λ, and therefore cannot distinguish states.
- Quantum-enhanced experiments: Two-copy invariant models span operators built from per-qubit identity and SWAP terms, providing exponentially many choices for eO(θ).The space C^(2)(G) is spanned by 2^n elements.
- Quantum-enhanced experiments: Suitable two-copy choices recover several multipartite entanglement measures, including concurrence, Concentratable Entanglement, and n-tangle families.
- Quantum-enhanced experiments: Because selected invariant operators produce different outputs across classes, the resulting models can perfectly classify the multipartite entanglement dataset.
V. DISCRETE GROUP-INVARIANT MODELS
The framework extends to discrete symmetries by encoding graph-isomorphism classes into quantum states with permutation symmetry. Symmetry-constrained one-copy models provide a polynomially growing solution space for classification.
- Dataset construction: The graph-isomorphism dataset contains quantum states encoding graphs isomorphic to one of two fixed non-isomorphic reference graphs.
- Permutation symmetry: Permuting the qubits conjugates each state into one whose interaction graph is isomorphic and therefore has the same label.
- Permutation symmetry: The dataset’s symmetry group is the symmetric group S_n, acting through permutations of the encoded graph states.
- Invariant models: Conventional one-copy models can be constructed from operators in span(A^⊗n) that are invariant under S_n.
- Invariant models: The invariant solution manifold grows polynomially with n, following the Tetrahedral numbers.
- Scope: The framework can also specialize to permutation subgroups such as reflections and translations in condensed-matter and quantum-chemistry settings.
VI. EQUIVARIANT QUANTUM NEURAL NETWORKS
The section develops a modular route to group-invariant QML models by composing equivariant intermediate maps with an invariant final measurement. This broadens the framework beyond a restricted hypothesis class while leaving quantum-neural-network parameterization unspecified.
- Scope: The framework specifies the form of a G-invariant model but does not prescribe how to parameterize U(θ) or choose the measurement operator.Those design choices remain outside the framework’s prescription.
- Constructive decomposition: General QML models can be decomposed into M parameterized maps representing quantum neural networks, measurements, post-processing, or data encoding.The maps are composed sequentially, with each map using a subset of the model parameters.
- Equivariance: Equivariance relaxes map-level invariance: group-shifting an input produces a correspondingly group-shifted output.Intermediate equivariant maps propagate the group action to a final invariant map, ensuring invariance of the composed model.
- Equivariance: A G-invariant model can therefore use equivariant intermediate maps and a G-invariant final measurement, rather than imposing invariance on every component.This construction preserves the symmetry while permitting more flexible internal transformations.
- Quantum neural networks: For unitary QNNs, equivariance is obtained when U(θ) belongs to the k-th commutant C^(k)(G), with graph classification providing an S_n-equivariant example.The relevant condition applies to a QNN or its individual layers acting on k copies.
- Constructive decomposition: The decomposition is modular because equivariant maps can be identified and reused across models, while additional encoding and post-processing steps fit naturally into the same framework.The approach also covers models more general than those in Hypothesis Class 1.
VII. CONCLUSIONS
The conclusions show that symmetry-informed QML models recover or extend methods across purity, time-reversal, entanglement, and graph-isomorphism tasks. They position the framework as an early step toward geometric QML while identifying hardware, model-class, task-scope, and quantum-advantage boundaries.
- Contributions: The framework derives conditions for G-invariance from representation theory and recovers literature algorithms across several supervised QML tasks.The showcased tasks include purity, time-reversal dynamics, multipartite entanglement, and graph isomorphism.
- Purity classification: In the purity task, conventional experiments cannot provide a G-invariant classifier, whereas coherently accessing two copies enables classification through the SWAP expectation value.The symmetry group is unitary because unitary transformations preserve spectral properties.
- Time-reversal classification: For time-reversal classification, both conventional and quantum-enhanced models distinguish the states, recovering Bell-basis measurement and, with access to preparing unitaries, an exponentially lower experiment count.The Bell measurement operator arises from the Brauer algebra, while the exponential experiment reduction is associated with the quantum-enhanced setting.
- Entanglement classification: For multipartite entanglement, existing entanglement measures are special cases of the framework’s G-invariant model family, while additional-copy local permutation measurements are conjectured to yield new measures.The relevant symmetry is the n-fold direct product of the local unitary group.
- Graph isomorphism: For graph isomorphism, the framework extends beyond continuous Lie groups and identifies S_n-invariant models that classify graph-encoded quantum states.The task distinguishes whether a state belongs to one isomorphism class or another.
- Outlook: The work is presented as an early step toward a general theory of QML models with sharp geometric priors based on dataset symmetries.The authors envision a broader field of geometric quantum machine learning.
- Open questions: Most QNNs are not equivariant, and developing architectures compatible with near-term hardware remains an open direction involving circuit depth and connectivity requirements.Existing equivariant examples are described as exceptions rather than the norm.
- Scope and limitations: The study focuses mainly on binary supervised classification and does not cover more general model classes, while G-invariance alone does not determine quantum computational advantage.The authors point toward regression, unsupervised learning, post-processing, randomized measurements, and favorable scaling as extensions or unresolved questions.
Appendix A: Hermitian part of the commutant
The appendix proves that the k-th commutant is closed under Hermitian conjugation and uses this closure to construct Hermitian symmetry-compatible operators. It then applies these facts to conditions for invariant models involving Lie-algebra components and tensor-product observables.
- Hermitian closure: If A belongs to the k-th commutant C^(k)(G), then its Hermitian conjugate A† also belongs to C^(k)(G).The proof takes the conjugate transpose of the commutation relation defining the commutant.
- Hermitian construction: Any non-Hermitian commutant element can be converted into Hermitian elements A + A† and i(A − A†), both remaining in C^(k)(G).These constructions provide Hermitian operators suitable for use in invariant-model conditions.
- Invariant-model conditions: For a Lie-group model with i1 belonging to the Lie algebra, G-invariance follows when the state lies in the Lie algebra and the effective observable lies in the span of tensor products A_j ⊗ A_j.The operators A_j are specified relative to the j-th copy and the remaining copies.
- Invariant-model conditions: The proof treats tensor-product observables term by term and uses closure of the Lie algebra under the associated group action to preserve the relevant support condition.The expectation-value argument is extended from one tensor-product term to linear combinations.
- Proof structure: The resulting appendix establishes the algebraic closure and trace-based identities used to complete the corresponding invariant-model proofs.The final steps invoke the definitions of the commutant and group invariance.
Appendix C: Proof of Propositions 3 and 4
The appendix analyzes invariant models for datasets with distinct class-dependent symmetry groups and relates purity classification to ancilla-based implementations of the SWAP operator.
- Proof setting: Propositions 3 and 4 address classification when the two labels have different symmetry groups, G0 and G1.Each group has its own higher-order symmetries; Lie-group cases also involve Lie algebras and orthogonal complements.
- Proposition 3: A model is G0-invariant but not necessarily G1-invariant when its effective observable belongs to G0 but not G1.If the effective observable belongs to both groups’ relevant commutants, the model is invariant under both groups.
- Proposition 4: For both class-dependent invariances, the effective observable must satisfy the stated span and Lie-algebra conditions involving ρ and the operators Aj.The asymmetric case changes which Lie algebra and orthogonal-complement conditions are imposed for each class.
- Purity classification: Purity classification requires second-order structure: k = 1 cannot classify, whereas k = 2 can classify using the SWAP operator up to additive and multiplicative constants.The appendix then motivates evaluating SWAP expectations with an ancillary qubit and two copies of ρ.
- Ancilla-based models: Ancilla-based models use a 2n + 1-qubit quantum neural network and measure an ancilla observable whose conjugated form commutes with 11 ⊗V ⊗V.The resulting condition is eOA ∈ iu(2) ⊗C(2)(G), with the relevant symmetry action on the two copies of the state.
- Ancilla-based models: Theorem 9 guarantees invariant models that perfectly classify the purity dataset when eOA(θ) = A ⊗S, with A|0⟩ = |0⟩ and S containing a nonzero SWAP component.The choice eOA(θ) = Z ⊗SWAP recovers the operator measured by both the SWAP Test and the alternative ancilla-based algorithm.
- Circuit connection: The two ancilla circuits implement distinct unitaries but measure the same purity operator, A ⊗SWAP with A = Z.The circuit description identifies U1 as the canonical SWAP test and U2 as a CNOT-reduced machine-learning-derived algorithm.
- Classification error: When the class-1 output is fixed at c, classification is unambiguous if c lies outside [b1, b2]; otherwise some class-0 states can be misclassified.The misclassification probability is tied to P(c|0), and additive estimation errors extend the same analysis to a tolerance interval.
Appendix F: Concentration results for time-reversal datasets
The concentration analysis shows that, for Haar-random states in the time-reversal setting, model outputs can concentrate exponentially near the class boundary, making classification require exponentially many measurements.
- Motivation: Some time-reversal models require many experiment repetitions when the class-0 expectation approaches the class-1 value c.The difficulty arises when ⟨X⟩0 is approximately c, reducing the separation available for classification.
- Concentration bound: For purely imaginary Hermitian observables, the model variance scales as O(1/2^n), so Haar-random outputs concentrate exponentially around mean zero.The derivation uses the absence of identity support for such observables and assumes Tr(...) = 11 in the stated tensor-product case.
- Time-reversal dataset: For the relevant class-1 states, the model output is fixed at zero, while class-0 outputs also concentrate exponentially around c = 0.Consequently, an exponential number of shots would be needed for classification.
2. Quantum-enhanced experiments
This section specifies general QGNN layers and explains how parameter tying yields permutation-invariant QGCNN architectures for graph-structured quantum data.
- QGNN construction: A general QGNN applies Q Hamiltonian evolutions repeatedly P times, with variational parameters controlling graph-topology-dependent interactions.The Hamiltonians include terms acting on graph edges and nodes.
- Permutation invariance: Permutation invariance is enforced by tying parameters across vertex indices, removing those indices from the edge and node coefficients.The resulting generators use edge- and node-independent index sets.
- Permutation invariance: The resulting QGCNN generators are invariant under graph-preserving permutations of node indices and generalize Quantum Alternating Operator Ansatze.Invariance follows because the summation indices can be relabelled under those permutations.