Source-linked AI summary

Towards unsupervised representation learning for quantum data: quantum models with inference and generation

Robin Lorenz, Eric Brunner, Marcello Benedetti

arXiv:2609.00372v1quant-phcs.LG

TL;DR

Quantum technology may produce coherently available quantum states, but classical unsupervised representation learning relies on chain-rule factorisations that lack a universal quantum counterpart. The paper develops a state-over-time framework for joint visible-latent models with inference and generation, and characterises when such models exist. It finds a hierarchy of admissible states linked to PPT structure, with the Leifer-Spekkens construction supporting inference and generation for PPT model classes while permitting genuinely quantum correlations.

  • Problem

    Quantum data may be available as coherent quantum states, yet quantum theory lacks a universal standard factorisation analogous to the classical chain rule for connecting inference, generation, and joint models.

  • Method

    The paper defines joint visible-latent quantum models and inference or generation through channel-state factorisations induced by state-over-time maps, including training and data-extension notions.

  • Results

    The paper characterises ambiguous states for representative state-over-time maps, establishes a hierarchy related to PPT states, and constructs Leifer-Spekkens models with inference and generation.

  • Takeaways & Limitations

    The Leifer-Spekkens construction supports quantum models with inference and generation for PPT states, including genuinely quantum visible-latent correlations.

  • Takeaways & Limitations

    Quantum extended inference lacks a general posterior-like construction that always saturates the data-processing inequality, and the proposed direct channel approach changes the desired classical analogy.

Abstract

from arXiv · show

With quantum sensors, simulators and networks emerging, a future of quantum technology may produce quantum states as data---that is, coherently rather than as classical measurement records---thus motivating the study of suitable quantum generalisations of modern machine learning, including the automated, unsupervised extraction of useful representations. Two ingredients are central to the latter: inference, mapping observations to latent representations, and generation, mapping latent states back to synthetic data. Both are related to each other and to joint distributions for training models by the chain-rule of classical probability theory. The fact that quantum states however lack such universal, standard factorisation property thus poses a challenge. Here we develop a conceptual and mathematical framework for unsupervised representation learning from quantum data. Models are joint quantum states over visible and latent systems; state-over-time maps provide a notion of factorisation into a marginal state and inference (generation) channel; models with inference (generation) are ambiguous states---states for which such factorisation obtains---subject to a further consistency condition on extended inference maps as data extension. These stipulations are restrictive: we show that non-trivial models must feature non-linear such maps to the extended space. For three representative state-over-time maps, we completely characterise the ambiguous states, uncovering a hierarchy tied to the positive-partial-transpose (PPT) criterion from entanglement theory. Notably, the Leifer-Spekkens construction supports inference and generation exactly for model classes of PPT states, thus allowing genuinely quantum visible-latent correlations. We also formulate quantum counterparts of exact and approximate inference training, explore weaker notions of data extension and sketch a future research programme.

1 Introduction

The paper motivates unsupervised representation learning for coherently available quantum data and develops a framework based on inference, generation, and quantum state factorisation.

  • Motivation: Unsupervised representation learning aims to automate extraction of useful features from high-dimensional, multimodal, and heterogeneous data.Classical applications include classification, clustering, anomaly detection, and next-token prediction.
  • Motivation: Quantum sensors, networks, simulators, and memories may produce quantum states as data rather than classical measurement records.This motivates learning tasks directly from coherently available quantum data.
  • Classical ingredients: Inference maps visible observations to latent representations, while generation maps latent states to synthetic data.Classically, both maps are stochastic and compatible with a joint distribution used in training.
  • Quantum challenge: Quantum theory lacks an obvious standard analogue of the classical probability chain rule, making factorisation into inference or generation maps nontrivial.The paper relates this challenge to the no-cloning theorem.
  • Related work: Existing quantum machine-learning approaches generally do not combine a joint visible-latent quantum model with compatible inference and generation factorisations.Quantum autoencoders and related models use directly parametrised operations rather than factorisations of a common joint model.
  • Contribution: The paper develops a conceptual and theoretical framework using state-over-time maps to define quantum factorisation, models, and training notions.It derives characterisation and no-go theorems for model classes based on representative state-over-time maps.

2 Preliminaries

The preliminaries review classical latent-variable representation learning and explain why its factorisations, inference procedures, and training methods do not transfer straightforwardly to quantum systems.

  • Classical setup: Classical representation-learning models represent data with visible variables X and latent variables Y.Datasets consist of independently sampled realisations of X, while Y supplies a latent representation.
  • Classical factorisation: The classical chain rule factorises any joint distribution into a marginal and conditional in either inference or generation direction.Conversely, a marginal distribution and conditional stochastic map define a joint distribution.
  • Model components: A classical model’s joint distribution induces an encoder Pθ(Y|X) and decoder Pθ(X|Y), with generative models commonly specified by a decoder and latent prior.The encoder can be recovered from the model posterior.
  • Training: Training commonly minimises a divergence between the data distribution and model marginal, often using stochastic gradient descent.The KL divergence is used concretely, while the framework extends to broader divergence families.
  • Inference: Exact inference extends data with the model posterior and can avoid explicitly computing high-dimensional latent-space integrals.Approximate inference instead uses an independently parametrised encoder, trading accuracy for efficiency through bias and variance.
  • Quantum challenge: Quantum generalisation is difficult because valid joint extensions and posterior-like optimal inference are not generally available.Even when recoverability can saturate the data-processing inequality for some states, the required extensions may be unavailable or trivial.

3 The framework

The framework generalises classical representation-learning components to quantum data by modelling visible–latent joint states and separating data extension from extended inference. It identifies stringent consistency requirements, including nonphysical extended maps and nontriviality constraints.

  • Framework: The framework generalises key classical representation-learning ideas to fully quantum models while remaining broad enough to study principled and practical model classes.The authors describe these as strong ideal notions in some respects but do not claim every instance is interesting or efficiently implementable.
  • Quantum data and models: A quantum data source is an ensemble of states queried independently, while the user generally does not know the source’s ensemble decomposition.The associated data state is the ensemble average, but individual queries may return different component states.
  • Quantum data and models: Models are parametrised joint states over observed system X and latent system Y, with the marginal on X required to represent the quantum data.Classical parametrised joint distributions appear as a special case after encoding them in a product basis.
  • Quantum data extension: A physical CPTP data-extension map satisfying the required marginal condition is necessarily trivial, appending a fixed state on Y independently of the input.Proposition 4 formalises this as E = I_X ⊗ η.
  • Quantum models with inference and generation: Quantum factorisation is not straightforward: inference channels must be related to joint states through a suitable state-over-time construction.The framework therefore separates data extension from extended inference, whose requirements are strictly stronger.
  • Quantum models with inference and generation: The framework’s extended inference and generation maps cannot generally be expected to be CPTP, because directly imposing positivity and the marginal condition yields useless discard-and-prepare inference.The same issue applies analogously in the generative direction.
  • Towards representation learning: Approximate inference training jointly optimises model parameters and a family of data-extension maps rather than requiring extension to be induced by the current model.This extends the framework beyond exact model-consistent inference.

4 Examples of state-over-time maps

The paper studies several state-over-time maps built from channel-state representations and star products, comparing their algebraic, positivity, and linearity properties. These maps induce distinct classes of ambiguous states used later to analyse quantum inference and generation.

  • Channel-state representations: The construction uses Jamiołkowski and Choi representations to associate operators with channels, with the two representations related by partial transposition.The Jamiołkowski representation is basis-independent, whereas the Choi representation depends on a chosen basis.
  • Examples of state-over-time maps: The selected state-over-time maps are based on ordinary operator, reverse-order operator, Jordan, and Leifer–Spekkens star products.Each map combines the channel representation with an operator on the input system.
  • Algebraic properties: Ordinary and reverse-order operator products yield linear trace-preserving maps but generally fail to preserve hermiticity.The operator product is bilinear, while the reverse-order product uses the opposite multiplication order.
  • Algebraic properties: The Jordan-product map is linear, trace-preserving, and hermiticity-preserving, but generally not positive.Thus it can preserve Hermitian structure without mapping every positive input to a positive output.
  • Algebraic properties: The Leifer–Spekkens map is nonlinear in its second argument, hermiticity-preserving, and trace-preserving; its positivity is equivalent to positivity of DJ(E).It is linear in the channel representation’s first argument.
  • Ambiguous states: Each selected state-over-time map defines a corresponding class of ambiguous states, including LS-, JP-, OP-, and PO-ambiguous states.The paper studies these classes in the factorisation direction determined by the visible and latent systems.

5 Non-trivial models need non-linear maps

The framework requires extended inference or generation maps to be non-linear for models to be non-trivial. Linear positive extensions collapse to trivial product-state models.

  • For any state-over-time map, product states are ambiguous and therefore provide the universally available trivial models.
  • Triviality identifies product states, discard-and-prepare channels, and extensions that factor into IX and a fixed state.
  • If an extended inference or generation map is linear and positive, the corresponding channel and state are necessarily trivial.
  • Non-trivial models therefore require non-linear extended maps; this rules out JP, OP, and PO as interesting constructions while leaving LS as a candidate.
  • A linear positive data-extension map must factor as the identity on X tensored with a fixed state on Y.

6 Ambiguous states

The paper characterises ambiguity classes for several state-over-time maps and relates them through PPT-based inclusions. The Leifer-Spekkens class is exactly PPT, while Jordan-product ambiguity extends beyond PPT.

  • Preliminaries: PPT states correspond, up to an X basis choice and normalization, to completely positive maps with positive Jamiołkowski representations.
  • 6.2 Operator-product ambiguity: OP ambiguity is exactly characterised by the CX-PPT condition: the state is PPT and commutes with its visible marginal.
  • 6.3 Leifer-Spekkens ambiguity: Leifer-Spekkens ambiguity is captured precisely by the PPT condition, with ambiguity-realising channels essentially unique on the marginal’s support.
  • 6.4 Jordan-product-based ambiguity: Jordan-product ambiguity is exactly characterised by DX-PPT states, a class strictly larger than PPT and containing non-PPT states.
  • 6.5 Summary: The classes form a strict hierarchy, and CX-PPT states can include non-separable PPT states with genuinely quantum correlations.
  • 6.5 Summary: Maximally entangled states are not DX-PPT and consequently are not ambiguous under any of the studied notions.

7 The LS-construction: models with inference and generation

The LS construction yields non-trivial quantum models with inference and generation, but only when extended maps are non-linear. Its valid model classes are characterised exactly by the PPT condition.

  • Non-linearity: Non-trivial LS-based models require non-linear extended inference or generation maps; non-linearity is necessary but not sufficient.The operator-product, product-operator and Jordan-product maps are ruled out for non-trivial models, while the LS construction meets the requirement.
  • PPT characterisation: For the LS construction, a model class supports inference and generation if and only if every model state is PPT.The result applies symmetrically to the visible and latent systems.
  • PPT characterisation: Every PPT state can be factorised in both directions into inference and generation channels under the LS state-over-time map.The induced extended maps are positive and trace-preserving, so they map states to states.
  • Consequences: The LS-based framework establishes that the proposed stipulations are non-vacuous by identifying model classes with inference and generation.The paper presents further study of LS models and the broader landscape of state-over-time maps as next steps.

8 Data extension without inference and generation

The paper distinguishes generalised data extension from extended inference and generation, allowing weaker model-of-data notions that need not arise from a state-over-time map. These weaker notions broaden practical flexibility but are not compatible with SOT-based inference training.

  • Motivation: Quantum data extension is distinguished from extended inference because a valid extension need not be induced by an inference map.The distinction is stronger than in classical models, where the concepts are not normally separated.
  • Generalised data extension: A generalised extension map sends states on X to joint states on X ⊗ Y while allowing corrected measurement statistics to reproduce the data expectations.A model is a model of data when it is the image of the data state under such a map.
  • Examples: Channels built from the B+ and B− constructions provide examples of generalised data extension maps.These channels use the SWAP operator and can be combined to implement virtual broadcasting.
  • Limitation: Generalised data extension is not compatible with training an SOT-based model with inference or generation.The incompatibility follows because the extension map is not itself an SOT map and need not output a representation of the input observable.
  • Limitation: The usefulness of weaker extension approaches depends on the success and practical costs of training models with inference and generation.This leaves their practical value conditional on the focused framework’s trainability.

9 Discussion

The discussion consolidates the framework’s structural results: non-trivial models require non-linear extensions, and LS models are exactly the PPT states. It also identifies open questions about trainability, expressivity, and broader constructions.

  • Framework: The framework models quantum data using joint visible-latent states and state-over-time factorisations for inference and generation.It distinguishes extended inference from the weaker notion of data extension and supports exact and approximate training setups.
  • Structural result: A linear, positive extended map that exactly preserves its input marginal must be trivial, so non-trivial models require non-linear extensions.The operational inference and generation maps themselves remain ordinary quantum channels.
  • Structural result: PPT is the precise boundary for valid bidirectional LS factorisations and includes PPT-entangled states while excluding non-PPT states.Thus the LS construction permits genuinely quantum visible-latent correlations, not only separable ones.
  • Comparison: The Jordan-product construction admits more ambiguous states than LS, but its linearity prevents non-trivial models under the positivity requirement.Ambiguity alone is therefore necessary but not sufficient for a globally valid learning model.
  • Limitations: The structural results do not establish efficient trainability, and the effects of PPT on inductive bias and expressivity remain empirical questions.The paper also makes no claim of quantum advantage in the usual or rigorous sense.
  • Future work: Future work includes approximate and robust formulations, tractable expressive PPT parametrisations, experiments, and broader state-over-time maps.A more radical option would relax positivity, potentially restoring linearity but requiring meaningful learning procedures for non-physical operators.

A.1 Training and the gradients of quantum relative entropy

The paper sketches gradient-based training with quantum relative entropy while emphasising that the relevant gradients are non-linear functionals of quantum states. Multi-copy estimation strategies may help, but tractability remains model-dependent and open.

  • Training objective: Training is formulated as minimising an objective such as quantum relative entropy over model and data-extension parameters.For practical reasons, optimisation would normally alternate between parameter sets rather than update both simultaneously.
  • Gradient structure: The framework’s gradients are inevitably non-linear in the given states because extension maps, generation maps and objective arguments introduce non-linearity.This applies both to gradients with respect to model parameters and to gradients with respect to extension parameters.
  • Gradient estimation: The paper rewrites relevant non-linear functionals using observables and maps, with Monte Carlo approximation providing a route to estimate integral expressions.The resulting gradient can become an expectation-value estimation problem.
  • Gradient estimation: Polynomial state functionals can be estimated as expectation values of observables on multiple independent copies of the state.Controlled-SWAP and cyclic-permutation tests are cited as basic examples of this strategy.
  • Limitation: Whether practical implementations are tractable depends on the chosen parametrised model class and is left to future work.The paper specifically identifies LS and other suitable state-over-time constructions as model-class choices.

A.2 Example of a non-PPT state that is DX-PPT

A concrete 3×2 bipartite example shows that DX-PPT strictly contains PPT by including a state with distillable entanglement.

  • The example uses dX = 3 and dY = 2, defining ρ from real and imaginary components.
  • The eigenvalues of ρ, TX(ρ), and TX(KρX(ρ)) are evaluated to test state positivity and the DX-PPT condition.
  • −0.01634675 is an eigenvalue of TX(ρ), establishing that ρ is not PPT.
  • The example is non-separable beyond bound entanglement because its entanglement is distillable.
  • A brute-force random search found approximately 2.2% such examples for dX = 3, dY = 2 and 0.1% for dX = 4, dY = 2.The search generated 1000 random states in each setting.

A.3 On restricted JP-based inference and virtual broadcasting

The JP-based construction generally fails to provide positive extended inference for nontrivial states and channels, motivating restricted-input alternatives and virtual broadcasting approximations.

  • For the JP construction, positivity of the extended map occurs only when the state and channel are trivial in the respective senses defined by the framework.
  • The JP map is a linear approximation to the LS map near the maximally mixed state and may yield physical statistics for highly mixed inputs.
  • Inference is considered sound only when the extended JP map also produces a valid joint state, preserving consistency with model training.
  • Restricted JP-based inference is proposed as a possible alternative for selected non-separable states, but the relevant input-state sets remain to be characterised.
  • The virtual broadcasting map B is linear, hermiticity-preserving, and trace-preserving, but is generally not positive or physically implementable.
  • Keeping only the physical channel B+ sacrifices an extended inference map while allowing arbitrary state inputs and producing output states.
  • Allowing both summands for arbitrary inputs would leave the framework of quantum-state models because the resulting objects need not be positive operators.

B.2 Proofs Sec. 5

These proofs establish structural properties of SOT maps, including product-state ambiguity and positivity constraints for the associated constructions.

  • For any SOT map, every product state α ⊗β belongs to the corresponding ambiguous-state set.
  • The proof uses simultaneous diagonalisation in a product basis to identify the constructed state with α ⊗ρY.
  • Membership of a marginal channel in the ambiguous set is equivalent to the SOT construction acting as tensoring with the marginal state.
  • For positive operators A and B, positivity of A^1/2BA^1/2 is equivalent to positivity on the support of A.
  • The rank-1 argument uses Schmidt decomposition and support inclusion, then extends to general positive states through eigendecomposition.

B.4 Proofs Sec. 6.2

The proofs characterise Sec. 6.2 state classes through commutation, partial-transpose positivity, and Choi-operator constructions of compatible channels.

  • For OP-based models, a joint state represented as DJ(E)ρX must satisfy commutation between DJ(E) and ρX together with a support-positivity condition.
  • The proof derives CX-PPT structure by combining commutation with partial-transpose positivity in an eigenbasis of ρX.
  • A CX-PPT state yields a compatible CP map by constructing a positive Choi operator and extending its action outside the support of ρX.
  • The channel representation is unchanged by modifications outside the support of the marginal, allowing extensions E′ with the same restricted action.
  • The LS-based proof obtains the channel and Choi-state relations by multiplying the factorisation with the inverse square root of the marginal on its support.

B.6 Proofs Sec. 6.4

These proofs characterise states satisfying the Jamiołkowski-product condition through positivity, trace, and complete-positivity properties. They also relate the Leifer–Spekkens and Jamiołkowski constructions through explicit map identities and star products.

  • A state satisfies the defining condition when ρ = Fρ ⋆JP ρX, Fρ ⪰ 0, and trY(Fρ) = πX.These are presented as the three characterising properties of the state.
  • The proof identifies Fρ as the unique solution of a Lyapunov equation, with uniqueness restricted to the support of ρX.The associated map is composed with ΠX, and Hurwitz stability supplies the relevant uniqueness condition.
  • The proof of complete positivity for Kη uses an eigendecomposition of ρX and the hyperbolic secant's Fourier-transform fixed-point property.The resulting Kη is a trace-non-increasing completely positive map.
  • For the Leifer–Spekkens construction, positivity of the corresponding Jamiołkowski representation is established through the relation C ⊙LS η = η1/2DJ(C)η1/2.The argument uses positivity of DJ(C) and the definition of the g AC construction.
  • The two ambiguity constructions are connected by star-product identities between their Jamiołkowski representations and the maps KρX and DJ(ELS).The text states that DJ(ELS) = KρX(ρ) ⋆JP ρX^-1, while the other construction is related through Prop. 23(2).
Loading 2609.00372v1…