Source-linked AI summary

"Train classical, deploy quantum" requires rethinking generalization

Snehal Raj, Natansh Mathur, Alejandro Perdomo-Ortiz

arXiv:2608.31117v1quant-phcs.LG

TL;DR

The paper asks whether classically evaluated training losses in train-classical, deploy-quantum workflows yield generators that generalize beyond training statistics. It benchmarks classical and quantum models and analyzes moment losses theoretically, finding that moment-trained models generally generalize worse than likelihood-trained ones. Thus, converged moment-matching loss is not a reliable measure of generalization.

  • Problem

    It remains unclear whether minimizing classically evaluated moment-matching objectives produces generalizing generators rather than models reproducing training statistics.

  • Method

    The authors benchmark thirteen classical and quantum models by direct sampling on two datasets up to 30 qubits and analyze moment-loss minimizers theoretically.

  • Results

    Moment-trained models generally generalize worse than likelihood-trained models, and small moment losses do not track coverage or forward KL.

  • Takeaways & Limitations

    A converged moment-matching loss is not a reliable sign that a train-classical, deploy-quantum model is a useful generator.

  • Takeaways & Limitations

    The study covers specific model families and losses on datasets whose valid sets are known and cheap to verify; real deployments may lack a reliable validity oracle.

Abstract

from arXiv · show

Generative models have become central across science and industry, from image and text synthesis to the design of molecules and materials. Quantum generative models are considered one of the most promising applications for quantum computers, since a quantum circuit naturally produces samples from the distribution it encodes, and for suitable circuits that distribution is believed to be hard for any classical computer to reproduce. A leading strategy trains these models on a classical computer and reserves the quantum device for generating samples at deployment. This is possible when the training loss can be evaluated on a classical computer. A prime example is the maximum mean discrepancy (MMD$^2$), a moment-matching loss that compares the model and the data through their Pauli-$Z$ correlations. Research so far has asked whether such models can be trained and whether their sampling is hard; whether minimizing such an objective yields a model that generalizes, rather than one that merely reproduces the training statistics, remains poorly understood. We benchmark a broad set of quantum and classical generative models by direct sampling and show that models trained with a moment-matching loss generally show worse generalization than the likelihood-trained models. We show this on two application-inspired datasets: first a cardinality-constrained dataset at up to $30$ qubits and second a dataset of genomic single-nucleotide variants, whose valid set is the observed data. These results indicate that a converged moment-matching loss is not a reliable measure of generalization, and that train-classical, deploy-quantum workflows will need approaches that target generalization directly, leaving open whether better training objectives suffice or whether the model architectures themselves must change.

I. INTRODUCTION

The paper reframes train-classical, deploy-quantum workflows around whether low classical training losses produce useful generators that generalize. It combines benchmarking and theory to show that moment-matching objectives can fit training statistics without reliably indicating deployment generalization.

  • I. INTRODUCTION: Train-classical, deploy-quantum workflows evaluate losses classically and reserve quantum hardware for deployment sampling.This avoids the expense of hardware-based gradient estimation while relying on classically tractable quantities.
  • I. INTRODUCTION: Prior work emphasized optimizing classical losses and reproducing deployed models, while paying less attention to producing unseen, valid samples.The paper distinguishes sampling hardness from usefulness as a generator.
  • I. INTRODUCTION: The benchmark tests whether low negative log-likelihood, fixed-order correlator, and full-kernel MMD2 objectives correspond to low forward KL.The figure frames training objectives as quantities evaluated without sampling the deployed model.
  • I. INTRODUCTION: Thirteen classical and quantum models are benchmarked up to 30 qubits on cardinality-constrained and genomic-variant datasets using direct sampling.Models are compared through training loss, coverage, and forward Kullback-Leibler divergence.
  • I. INTRODUCTION: Moment-matching loss correlates poorly with coverage and forward KL, while those two generalization measures agree more closely.Moment-trained models generally generalize worse than likelihood-trained models.

II. BACKGROUND AND RELATED WORK

The related work motivates classical training for quantum generative models and separates trainability or simulability from generalization. It positions coverage and forward KL as measures that can distinguish memorization from useful generation.

  • II. BACKGROUND AND RELATED WORK: Quantum circuit Born machines encode distributions in measurement outcomes, with sampling native to hardware and some circuit families believed classically hard to sample.The section notes that learnability can nevertheless vary substantially across circuit families.
  • II. BACKGROUND AND RELATED WORK: Train-classical, deploy-quantum methods use classically computable Pauli-Z correlators to train quantum models without hardware-based sampling during optimization.Prior work demonstrated IQP Born-machine training with MMD2 at up to a thousand qubits.
  • II. BACKGROUND AND RELATED WORK: Earlier analyses primarily studied whether moment-matching losses are optimizable or whether quantum models have classical surrogates.These questions include barren plateaus, data-dependent initialization, and correlator-based descriptions.
  • II. BACKGROUND AND RELATED WORK: Training-distribution agreement alone cannot distinguish memorization from generalization, so the paper uses forward KL to the known target and coverage of unseen valid strings.Coverage is evaluated alongside fidelity when architectures do not confine support to the valid set.
  • II. BACKGROUND AND RELATED WORK: The tensor-network Born machine is likelihood-trained, whereas the generative moment-matching network uses MMD2, enabling comparison of loss effects apart from the quantum platform.Both are included among the evaluated classical models.

III. FRAMEWORK

The framework defines TCDQ losses as moment-based surrogates and evaluates generalization through forward KL and sample-based metrics. It also shows that low-order moment matching can admit exact minimizers with poor support coverage.

  • III. FRAMEWORK: A generative task uses a target distribution over binary strings, whose support S contains the valid strings and whose unseen valid set is U = S \ T.Coverage concerns the fraction of U receiving positive model probability.
  • III. FRAMEWORK: In TCDQ workflows, optimization quantities are evaluated classically, while deployment samples come from the quantum circuit.The model probabilities themselves may not be efficiently accessible even when selected correlators are classically computable.
  • III. FRAMEWORK: The three loss objects are the population discrepancy, empirical training loss, and finite-budget estimator: L∗(q), LT(q), and bLT,B(q).A small finite-budget estimate only indicates estimator accuracy and empirical optimization, not closeness to the target.
  • III. FRAMEWORK: A moment loss of capacity d depends on d bounded observable expectations and is minimized when those expectations match the reference.The deployed losses use Pauli-Z strings and selected correlator sets, with capacity determined by their span.
  • III. FRAMEWORK: The full MMD compares every correlator and is characteristic at the population level, but its empirical minimizer is the training-data memorizer.Thus an empirical MMD value cannot replace sample-based evaluation.
  • III. FRAMEWORK: Forward KL is finite only when the model assigns probability to every valid string and is evaluated offline because deployment probabilities and unknown targets are inaccessible.Coverage and fidelity provide sample-based deployment measures under a stated sampling budget.
  • III. FRAMEWORK: A truncated moment loss can have exact minimizers with identical low-order correlators but support coverages of 1 and 0.07.The framework uses this construction to demonstrate that moment matching alone does not certify generalization.

IV. MOMENT LOSSES DO NOT CERTIFY GENERALIZATION

Moment losses of bounded capacity can have exact minimizers with extremely poor coverage, so achieving a low loss does not certify generalization at deployment. Although the full MMD has the target as its unique population minimizer, empirical minimization can still select a memorizer.

  • Theoretical result: Exact global minimizers of bounded-capacity moment losses can be supported on at most d+1 valid strings when d+1<|S|.The target distribution is also an exact minimizer, so the loss cannot distinguish the sparse solution from the target.
  • Scope: The sparse-minimizer theorem constrains distributions, while whether a circuit family realizes such a minimizer or optimization reaches it requires empirical assessment.The paper addresses these separate questions in its benchmark.
  • Theoretical result: On the cardinality task, fixed-order correlator losses have capacity d = D_L = O(N^L) and admit exact minimizers supported on at most D_L strings.These bounds apply to the deployed losses on the uniform cardinality target.
  • Constructive example: At N = 12 and L = 2, a distribution matching every correlator through order two achieved truncated loss ∼10^-28 but coverage 0.07 versus 1.00 for the target, with infinite forward KL.The sparse minimizer was constructed by linear programming on 66 strings.
  • Empirical training: The empirical MMD can be minimized by the training-set memorizer, which has zero empirical loss and zero coverage, so sample-based evaluation is necessary.This limitation concerns empirical and stochastic training objects, not the unique exact minimizer of the characteristic-kernel population MMD.

V. BENCHMARKING GENERALIZATION

The benchmark compares training loss with direct generalization measures across thirteen classical and quantum models. Models with nearly identical MMD^2 losses can differ substantially in coverage, supporting poor alignment between moment matching and deployment generalization.

  • Benchmark design: Thirteen classical and quantum models are evaluated by training MMD^2, coverage, and forward KL to test whether low loss tracks generalization.For small cardinality systems, the authors also compare model rankings under these scores.

A. Datasets, models, and protocol

The study benchmarks likelihood- and MMD-trained classical and quantum models on cardinality-constrained and genomic-variant datasets using free sampling. Coverage, fidelity, and forward KL provide deployment-oriented measures alongside training loss.

  • Datasets: The cardinality dataset contains even-length bitstrings with Hamming weight N/2, and the uniform target is evaluated at N = 16, 20, and 30.The valid-set sizes are 12,870, 184,756, and 1.55 × 10^8, respectively.
  • Datasets: The genomic dataset contains deduplicated observed single-nucleotide-variant sequences, with valid-set sizes 2,716 at N=16 and 3,980 at N=20.Its target is the empirical distribution, which is non-uniform and includes linkage between loci.
  • Models: The likelihood-trained group includes TNBM, transformers, GRU RNNs, and RBM, while the MMD-trained group includes IQP, magic FBM, and GMMN.The passive FBM is confined to the cardinality valid set by construction and therefore has F = 1.
  • Protocol: Training subsets use fractions ε of each dataset, with shared subsets across models at each seed and resampling across seeds.The study uses three seeds for cardinality and five for genomic variants.
  • Reported results: Table I reports median results for cardinality at N = 30 and genomic variants at N = 20, grouping models by likelihood- versus moment-trained families.It reports eval MMD^2 for both datasets, fidelity and normalized coverage for cardinality, and coverage for genomic variants.
  • Evaluation: Free sampling retains all Q generated samples; coverage measures distinct unseen valid samples, fidelity measures validity, and forward KL compares the trained model with the known target.A random-bit sampler supplies a no-learning coverage floor, while normalized coverage corrects for query-budget limits.
  • Capacity analysis: Increasing IQP gate count eighteenfold does not raise coverage above the random baseline, whereas TNBM coverage rises before over-fitting.Table II examines capacity at N = 16 and ε = 0.10.

B. Cardinality-constrained data

On the cardinality-constrained data, moment-matching loss and generalization decouple: models can attain low MMD2 while covering little of the unseen valid set. Likelihood training and structural constraints produce substantially better coverage, with the gap persisting and widening through 30 qubits.

  • Cardinality-constrained data: The rank-correlation panels compare the same thirteen models at N = 16 using median-over-seed scores and Spearman’s r.The columns pair coverage with forward KL, then each generalization measure with MMD2; the top row is cardinality data.
  • Cardinality-constrained data: Moment-trained deploy-quantum models reach low MMD2 yet fail to generate valid bitstrings efficiently, whereas likelihood-trained models achieve comparably low MMD2 while covering most of the valid set.The passive fermionic Born machine is an exception because its circuit emits only Hamming-weight-preserving strings, confining outputs to the valid set by construction.
  • Cardinality-constrained data: 0.09 and 0.41 of the unseen valid set are covered by IQP and tensor-network Born machines with MMD2 values within 40%, a 4.6× coverage gap.At N = 16, coverage and forward KL agree closely, while MMD2 correlates poorly with both.
  • Cardinality-constrained data: An order-of-magnitude reduction in IQP MMD2 leaves coverage near Crand, while likelihood-trained transformer coverage rises toward the finite-sample maximum C∗.The figure also shows classical transformer, RNN, and TNBM coverage increasing with parameter count, unlike IQP and magic FBM.
  • Cardinality-constrained data: Increasing IQP gate count 18× leaves coverage flat, while enlarging tensor-network and classical models improves coverage.The decoupling is not explained by kernel bandwidth, and the gap widens with qubit count to N = 30.

C. Genomic variant data

On genomic single-nucleotide-variant data, deploy-quantum models are again the weakest generators, while transformer, RNN, and tensor-network models achieve the highest coverage. The relationship between MMD2 and generalization is dataset-dependent, so generalization must be measured directly.

  • Genomic variant data: Transformer and RNN models reach the highest genomic coverage, with the tensor-network Born machine just behind, all near C = 0.14.These models also fail to reach a low MMD2 because the dataset has no fixed Hamming weight to exploit.
  • Genomic variant data: Deploy-quantum models are again the weakest generators, and the passive fermionic Born machine generalizes poorly despite performing well on the cardinality data.Its Hamming-weight bias fits the cardinality set but not the genomic one.
  • Genomic variant data: MMD2 correlates slightly better with loss and coverage on genomic data than on cardinality data, but its relationship to generalization remains dataset-dependent.The paper therefore evaluates generalization directly using coverage and forward KL rather than relying on MMD2.

VI. CONCLUSION

The paper combines direct sampling benchmarks with a theoretical analysis to test whether moment-matching losses support generalization. It finds that moment-trained models generalize worse than likelihood-trained models and concludes that generalization should be checked through sampling, while noting important scope limits.

  • VI. CONCLUSION: Thirteen classical and quantum generative models are sampled and scored on two application-inspired datasets at up to 30 qubits.The evaluation measures whether generated samples are new and valid.
  • VI. CONCLUSION: Moment-trained models generally generalize worse than likelihood-trained models, although moment matching still converges to small values that do not track coverage or forward KL.The fermionic Born machine is an exception because its circuit confines outputs to the valid set by construction.
  • VI. CONCLUSION: Exact matching of prescribed low-order correlators can yield either full coverage or exponentially small coverage because high-order correlators remain unconstrained.The theoretical result shows that two distributions can drive the moment loss to zero while differing drastically in coverage.
  • VI. CONCLUSION: The study argues that TCDQ workflows should assess generalization by directly sampling trained models rather than treating training loss as sufficient evidence.The paper identifies generalization as the key requirement for this evaluation.
  • VI. CONCLUSION: The benchmark covers specific model and loss families on datasets with known, cheaply checked valid sets; real deployments may require a reliable validity oracle.The authors identify replacing moment matching with a generalization-targeting loss as a direct next step.

Appendix A: Proof of Theorem 1

Theorem 1 shows that bounded-capacity moment losses can have exact global minimizers with drastically different support coverage and forward KL. The proof constructs sparse minimizers and full-support mixtures that preserve all matched moments while generalizing arbitrarily poorly.

  • Theorem 1: A moment loss depending on d bounded observables admits a global minimizer supported on at most d + 1 valid strings.When the constant function lies in the observable span, the support bound improves to d.
  • Theorem 1: The sparse minimizer can have coverage at most (d + 1)/|S| and infinite forward KL, while the full-support target is also an exact minimizer.Both distributions attain the same loss value despite their different generalization properties.
  • KL bound: Choosing η ≤ exp[−(R + 1/e)/α] makes the forward KL exceed any prescribed finite R, so exact loss convergence is compatible with finite or infinite KL.The construction uses the unseen target mass α and a binary-partition data-processing bound.
  • Proof construction: The support-reduction proof repeatedly preserves all d expectations while eliminating support points until at most d + 1 remain.A kernel vector with zero total mass supplies the direction for each support-shrinking step.
  • Evaluation metrics: Coverage is measured through free samples reaching unseen valid strings, alongside exploration, fidelity, rate, and normalized coverage metrics.The normalized coverage compares observed coverage with the finite-sample ideal C∗.

Appendix D: Additional cardinality-task results

Additional cardinality-task results show that moment-loss convergence remains decoupled from coverage across sizes, losses, and kernel choices. Likelihood-trained and structural models improve coverage, while IQP remains near the random baseline.

  • Training-loss decoupling: Every model’s MMD2 converges to the same numerical range at N = 16 and N = 20, while coverage separates them.The scatter compares models across deployment groups, with lower loss and higher coverage preferred.
  • Scaling to N = 30: The normalized-coverage gap between the best model and IQP grows from 0.28 at small N to 0.86 at N = 30.This widening occurs despite convergence of the models’ MMD2 values to the same numerical range.

Appendix F: Genomic-task diagnostics

Genomic-task diagnostics compare generated samples with reference population-genetics structure and cardinality-task training behavior. The supplied passages emphasize model-specific loss–coverage separation and probability mass placement relative to valid sectors.

  • Genomic diagnostics: Genomic diagnostics compare allele frequencies, linkage disequilibrium, Hamming-weight distributions, and PCA projections of generated and reference samples.The comparisons use ε = 20% and N = 20.
  • Training dynamics: At N = 16 and N = 20, MMD2 converges to a common numerical range while coverage separates across the five displayed training trajectories.The TNBM, transformer, and RNN increase coverage, whereas IQP and active and magic FBM remain near Crand.
  • Training dynamics: At N = 30, all models converge in MMD2 while normalized coverage separates; passive FBM, transformer, and RNN approach 1 and IQP remains near Crand.The corresponding trajectories use HW = 15, with |S| approximately 1.55 × 10^8.
  • Training dynamics: At N = 30, sampled IQP MMD2 stays approximately 1.4 × 10^-3 with coverage near Crand, while transformer NLL converges with coverage near the finite-sample maximum.The comparison shows distinct training behavior for the two models.
  • Sampling coverage: Coverage versus query budget increases toward C∗ for likelihood-trained and structural models but remains near Crand for moment-trained models at N = 16 and N = 20.The figure frames low coverage as persisting across the studied budgets rather than arising from finite sampling.
  • Probability-mass placement: On the N = 16, HW = 8 space, number-conserving FBM places all probability mass on the valid Hamming-weight slice, whereas moment-trained models assign substantial mass outside it.The grid covers the full 2^16 bitstring space, with intensity indicating sample probability on the valid slice.
  • Kernel robustness: IQP coverage remains 0.13 to 0.16 across kernel bandwidths from 0.1 to 16, near Crand, while likelihood-trained baselines reach approximately 0.5.Neither sharp nor broad bandwidth choices rescue IQP coverage.
Loading 2608.31117v1…