Source-linked AI summary

IQP Born Machines under Data-dependent and Agnostic Initialization Strategies

Sacha Lerch, Joseph Bowles, Ricard Puig, Erik Armengol, Zoë Holmes, Supanut Thanasilp

arXiv:2603.14576v1quant-phcs.LGstat.ML

TL;DR

IQP-QCBM training needs initialization strategies that avoid barren plateaus while preserving useful MMD-based learning. The paper analyzes full-angle, data-agnostic, and data-dependent starts, showing that target-aligned warm starts retain local trainability and can outperform agnostic alternatives, subject to data limitations.

  • Problem

    The paper addresses how initialization affects the trainability of MMD-trained IQP-QCBMs.

  • Method

    It derives loss-variance bounds for full-angle random, small-angle data-agnostic, and data-dependent initializations, extending the framework to nonlinear losses.

  • Results

    Full-angle all-to-all initialization exhibits exponential concentration, whereas suitable small-angle and data-dependent patches retain non-vanishing local variance.

  • Takeaways & Limitations

    Data-dependent initialization combines local variance with better target alignment and is reported to improve convergence and model fidelity over agnostic strategies.

  • Takeaways & Limitations

    Data-dependent initialization cannot guarantee success when available samples lack sufficient information about the target distribution’s structure.

Abstract

from arXiv · show

Quantum circuit Born machines based on instantaneous quantum polynomial-time (IQP) circuits are natural candidates for quantum generative modeling, both because of their probabilistic structure and because IQP sampling is provably classically hard in certain regimes. Recent proposals focus on training IQP-QCBMs using Maximum Mean Discrepancy (MMD) losses built from low-body Pauli-$Z$ correlators, but the effect of initialization on the resulting optimization landscape remains poorly understood. In this work, we address this by first proving that the MMD loss landscape suffers from barren plateaus for random full-angle-range initializations of IQP circuits. We then establish lower bounds on the loss variance for identity and an unbiased data-agnostic initialization. We then additionally consider a data-dependent initialization that is better aligned with the target distribution and, under suitable assumptions, yields provable gradients and generally converges quicker to a good minimum (as indicated by our training of circuits with 150 qubits on genomic data). Finally, as a by-product, the developed variance lower bound framework is applicable to a general class of non-linear losses, offering a broader toolset for analyzing warm-starts in quantum machine learning.

I. INTRODUCTION

IQP-QCBMs combine quantum sampling with classically tractable MMD training, but their trainability depends strongly on initialization. This work analyzes full-angle, data-agnostic, and data-dependent starts to explain barren plateaus and warm-start behavior.

  • Motivation: IQP-QCBMs offer classically evaluable low-body correlators for MMD training while retaining sampling tasks believed to be classically hard.This combination makes them a useful setting for studying trainability with classical resources while preserving potential quantum advantage.
  • Research gap: Full-angle random initialization is motivated by general training, but prior observations of exponential concentration lacked a theoretical explanation for IQP-QCBMs.The paper addresses this gap by analyzing loss-landscape variance across three initialization families.
  • Contributions: The study proves barren plateaus for full-angle random initialization, derives variance bounds for small-angle starts, and finds data-dependent starts better aligned with targets and faster in 150-qubit genomic-data training.The variance framework also extends to a broad class of nonlinear losses.
  • Quantum generative modeling: Quantum generative models use measured quantum states to produce samples from parametrized model distributions that approximate target distributions.QCBMs parameterize the sampling distribution through a quantum circuit and computational-basis measurement.
  • MMD loss: MMD implicitly compares model and target distributions by matching Pauli-Z correlators, with low-body weighting obtained for bandwidth σ ∈ Θ(√n).Matching all correlators would identify identical distributions, while the low-body regime emphasizes simpler correlators.

C. Exponential concentration

Exponential concentration makes typical random parameter points difficult to distinguish and train, motivating restricted initialization patches. The paper links patch variance to local curvature and examines data-agnostic and data-dependent warm starts.

  • C. Exponential concentration: Barren plateaus arise across variational and quantum machine-learning settings, including nonlinear losses and quantum generative models.The paper places concentration among broader scalability challenges such as limited expressivity and unfavorable optimization landscapes.
  • C. Exponential concentration: Exponential concentration causes loss variance to vanish with system size, making appreciable deviations require exponentially fine precision and measurement shots.Chebyshev’s inequality connects vanishing variance to the rarity of detectable deviations from the mean.
  • D. Initialization strategies: A globally flat landscape can still contain small regions with substantial gradients and favorable minima, motivating alternative initialization strategies.This observation underlies analyses of small-angle and other warm-start methods.
  • D. Initialization strategies: For nonlinear losses, polynomially large Hessian curvature at a patch center implies polynomially large local variance and non-exponentially vanishing gradients.The paper studies this curvature-variance link for initialization patches in generative modeling.
  • D. Initialization strategies: The paper compares identity, unbiased, and data-dependent marginal-matching initializations within restricted parameter patches.The identity and unbiased strategies are data-agnostic baselines, while the final strategy uses empirical target statistics.

III. RESULTS

The results separate global full-angle concentration from local curvature under restricted initialization. Full-angle all-to-all IQP circuits exhibit barren plateaus, whereas suitable small-angle and data-dependent starts retain trainable variance, with data-dependent initialization offering the strongest practical performance.

  • Small-angle initialization: Small-angle patches with non-vanishing center curvature have at least polynomially large correlator and MMD-loss variance, avoiding exponential concentration.The curvature decomposes into mismatch-driven and model-sensitivity contributions.
  • Data-agnostic initialization: Identity and unbiased data-agnostic patches both provide sufficient curvature, but identity starts from a delta distribution while the unbiased point gives a uniform distribution.The identity distribution is highly biased and can be detrimental, whereas the uniform distribution is a natural data-agnostic prior.
  • Data-dependent initialization: Data-dependent marginal matching guarantees inverse-polynomial local curvature under stated assumptions and empirically improves convergence speed and model fidelity over data-agnostic alternatives.The supplied passage states this as the paper’s data-dependent initialization result.
  • Full-angle random initialization: All-to-all connectivity drives every non-trivial correlator’s variance to vanish exponentially because its effective Pauli light cone spans the system.The correlator variance is governed by light-cone reach, which saturates at n under all-to-all connectivity.
  • Full-angle random initialization: Full-angle random initialization in all-to-all IQP circuits makes the MMD loss exponentially concentrate and ineffective for training at scale.The theorem’s bound is independent of the target distribution’s specific structure or correlations.

C. Evading barren plateaus with small-angle initializations

Small-angle initialization can create local regions with substantial curvature and non-exponentially vanishing loss variance, even when full-angle IQP landscapes exhibit concentration. The resulting variance framework extends beyond MMD to broader nonlinear losses, while requiring caution about concentrated loss components.

  • Local variance guarantees: Full-landscape concentration does not exclude smaller high-curvature regions where gradient signals remain accessible.This motivates analyzing local subregions rather than treating global barren plateaus as proof of untrainability.
  • Application to IQP-MMD: For MMD and IQP circuits, the general theorem’s derivative conditions are satisfied, establishing the framework’s applicability to this setting.The analysis provides the relevant constants for the IQP-MMD case.
  • Local variance guarantees: A curvature-based theorem guarantees a local patch with substantial loss variance for nonlinear losses satisfying bounded-derivative conditions.The result applies to arbitrary parametrized quantum processes, including IQP circuits and the MMD loss, with patch half-width r in O(1/poly(n)).
  • Caveat: Non-vanishing loss variance is not sufficient for successful training if the loss components estimated by the quantum computer themselves concentrate.The framework therefore indicates a possible escape from exponential concentration but does not independently guarantee trainability.
  • Curvature mechanisms: MMD curvature separates into data-mismatch and model-sensitivity mechanisms, respectively reflecting target disagreement and parameter-dependent output sensitivity.The mismatch term grows when model and target correlators are misaligned, while sensitivity is intrinsic to the architecture.

D. Data-agnostic initialization strategies

Identity and unbiased data-agnostic initializations provide local variance guarantees through different curvature mechanisms, while data-dependent initialization preserves sensitivity using target marginals. The formal guarantees rely on assumptions about the target distribution and do not by themselves establish superior training performance.

  • Identity initialization (Mismatch-driven): Identity initialization yields non-vanishing patch variance through target–model mismatch, but it need not provide a high-quality generative starting point.At θ=0, correlators are maximized, model sensitivities vanish, and curvature is driven by mismatch with the target.
  • Unbiased initialization (Sensitivity-driven): Unbiased initialization yields a non-vanishing optimization signal solely from IQP architectural sensitivity while its mismatch contribution vanishes.Setting single-qubit angles to π/4 and two-qubit angles to 0 produces the uniform distribution over computational-basis states.
  • Agnostic guarantees: Theorem 3 guarantees a non-vanishing signal in a local patch for identity and unbiased initializations under their stated conditions.The identity case assumes at least one system-size-independent single-bit marginal; the unbiased center uses π/4 single-qubit and zero two-qubit angles.
  • Data-dependent initialization: Data-dependent initialization matches single-body model correlators to target correlators and keeps two-qubit parameters at zero.Under the stated assumptions, target-modulated sensitivity remains bounded away from zero and Theorem 2 yields a local variance guarantee.
  • Data-dependent initialization: The data-dependent guarantee assumes approximately factorizable target correlators and a target that is not concentrated around one bit string.The authors describe these assumptions as technical sufficient conditions rather than strict requirements for practical use.
  • Data-dependent initialization: For low-body MMD, the data-dependent strategy guarantees non-exponentially vanishing loss functions, but numerical analysis is needed to compare training performance with agnostic strategies.Low-body MMD corresponds to bandwidth σ in Θ(√n), emphasizing low-body correlators.

F. Numerical studies

Numerical studies compare variance and training across agnostic and data-dependent initializations, finding broader useful variance regions and faster convergence for data-dependent strategies.

  • Loss variance: Data-independent initializations provide non-zero variance over noticeably narrower scale regions than data-dependent strategies.
  • Loss variance: The variance-maximizing initialization scale roughly grows polynomially with qubit count, with strictly larger values for data-dependent methods.
  • Training performance: The unbiased initialization remains trainable at linear and square-root scales but converges more slowly than the data-dependent approaches.
  • Training performance: Data-dependent strategies show rapid loss decay for s = 1/m and s = 1/√m, while identity initialization quickly stalls and fails to converge.The comparison uses 150-qubit IQP-QCBMs trained on genomic data.
  • Full-angle initialization: At full-angle scaling, most strategies fail, whereas covariance initialization can converge because correlated two-qubit parameters and scale rescaling alter the initialization range.

IV. DISCUSSION

The discussion separates global concentration from local curvature and emphasizes that data-dependent warm starts can improve alignment without guaranteeing optimization success. It also identifies limited sample information and post-initialization behavior as important boundaries.

  • Interpretation: Data-dependent initialization combines sufficient local loss variance with better target alignment at initialization, providing an advantage over agnostic strategies.
  • Limitations: The strategy cannot guarantee success when available samples lack enough statistical information to resolve the target distribution’s structure.Matching kth-order marginals does not strictly guarantee a signal for (k + 1)th-order marginals.
  • Limitations: For Haar-random target distributions, polynomially many samples yield correlators with no instance-specific information with exponentially high probability, leaving the MMD loss statistically insensitive to the target.
  • Open questions: Trainability should be studied beyond the first few steps because optimization may leave regions where initialization-based guarantees apply.

Appendix A: Notation table

The setup defines IQP correlators, interaction-graph neighborhoods, and effective Pauli light cones used to analyze full-angle initialization. The resulting variance depends on light-cone size and vanishes exponentially for all-to-all connectivity.

  • Notation and correlators: The model correlator z_A(θ) is the expectation of a Pauli-Z string supported on qubit subset A under the parametrized IQP circuit.Single-qubit correlators are written as z_j(θ).
  • Interaction graph: The external neighborhood N_E(A) contains complement qubits connected to at least one qubit in A through the interaction graph.It identifies the qubits added to the observable’s causal support.
  • Effective light cone: The effective Pauli light cone has size d_A = |A| + |N_E(A)| and counts qubits entering the Heisenberg-evolved operator support.This quantity controls correlator variance under full-angle random initialization.
  • Variance dependence: For full-angle initialization, Var_θ[z_A(θ)] = 2^-d_A, so highly connected graphs produce exponentially small correlator variance.The variance is governed purely by interaction-graph reach.
  • All-to-all topology: All-to-all connectivity gives d_A = n for every nontrivial A, causing uniform exponential concentration across correlators.Different correlators also have vanishing cross terms under full-angle random initialization.

2. Proof of Proposition 1: Exact correlator variance under full-angle random initialization

The proof derives exact correlator moments under full-angle random initialization by exploiting anti-commuting circuit generators and the interaction graph. It then transfers exponential correlator concentration to the full MMD loss for all-to-all IQP circuits.

  • Exact variance: Counting constrained bit strings yields Var_θ[z_A(θ)] = 2^-d_A, where d_A = |A| + |N_E(A)|.Qubits outside A and its external neighborhood remain free in the variance calculation.
  • Cross terms: Distinct nontrivial correlators have vanishing cross expectations under full-angle initialization, so their covariance contributions disappear.This cancellation does not require all-to-all topology when single-qubit rotations are present.
  • MMD consequence: For all-to-all IQP circuits with parameters uniform on [−π/2, π/2], the MMD loss itself concentrates exponentially and is generically untrainable across the full parameter landscape.The result follows by combining correlator concentration with vanishing cross terms in the MMD variance.

2. Polynomially large MMD variance for full-angle regime with restricted topology

Restricted interaction topologies can preserve inverse-polynomial MMD variance under full-angle initialization when low-body correlators remain sufficiently local. The analysis also motivates small initialization patches, where identity-centered correlators can remain trainable.

  • Sparse topology: For sparse d-dimensional lattice topologies with constant degree, low-body correlator variance can be inverse-polynomial rather than exponentially small.The paper gives 1D and 2D lattices as examples with K = 2 and K = 4.
  • Restricted topology: Under restricted K-regular topology, the MMD loss variance is inverse-polynomial when a low-body target correlator is nonzero and K = O(log(poly(n))).The bound uses low-body subsets with |A| = O(1) and the topology-dependent correlator variance.
  • MMD regime: Low-body MMD uses σ ∈ Θ(√n), which assigns polynomial-scale weight to small-support correlators while suppressing higher-body terms.Consequently, the variance guarantee directly targets low-body data or low-body approximations.
  • Identity initialization: Small patches around identity can retain non-vanishing correlator variance, showing that full-angle concentration depends on initialization patch size rather than IQP architecture alone.The identity-centered result also yields a tighter architecture-specific width characterization than generic curvature bounds.

2. The mixed fourth derivatives follow

The appendix extends curvature-based variance arguments from expectation-value losses to general nonlinear functions and specializes them to MMD losses in IQP circuits. Under derivative and curvature conditions, sufficiently small patches obtain polynomially large variance.

  • Patch guarantee: If the curvature is at least polynomially large, a patch with half-width r ∈ O(1/poly(n)) has a corresponding variance lower bound.The result links accessible local variance to a sufficiently small initialization neighborhood.
  • General nonlinear case: The framework derives a variance lower bound for general nonlinear loss functions from local curvature and bounded-derivative conditions.The proof uses multivariable variance decomposition and derivative bounds.
  • MMD specialization: For MMD losses in the stated IQP architecture, the general derivative conditions specialize to the circuit’s local curvature structure.The specialization covers single- and two-qubit gates.
  • Technical statement: The analysis introduces a proposition bounding the distance between a parametrized function’s local average and its fixed-point value.This proposition is identified as an original derivation central to the initialization-landscape analysis.

2. Proof of Theorem 2: Lower bound guarantee of an arbitrary non-linear loss and MMD loss with sufficient curvature

The proof derives a general variance lower bound for nonlinear losses and verifies the required derivative conditions for IQP MMD losses. The resulting framework yields an informative lower bound when local curvature is sufficiently large.

  • General nonlinear loss: The variance decomposition isolates a single-parameter contribution, providing a lower bound on the total variance through permutation-invariant parameter indexing.The proof reduces the analysis to Ēθ[Var_θα[L(θ)]], then bounds this term through successive variance, expectation, and curvature arguments.
  • General nonlinear loss: The lower-bound proof requires bounded higher derivatives and a sufficiently small patch radius satisfying the stated conditions on r and Δ.The argument uses derivative bounds, positivity of variance, and conditions ensuring the final difference remains non-negative.
  • IQP MMD specialization: For IQP MMD losses, the required derivative conditions hold with explicit constants a, γ, γ_j, and Δ determined from Pauli-correlator weights and gate anticommutation.The construction separately bounds correlator derivatives and sums the weights of terms that anticommute with one- and two-qubit generators.

Appendix H: Curvature mechanisms at initialization: identity, unbiased, and data-dependent centers

The curvature analysis distinguishes identity, unbiased, and data-dependent initialization by the mechanism producing local curvature. Identity relies on target-model mismatch, whereas unbiased and data-dependent centers can retain model-sensitivity contributions under their respective conditions.

  • Unbiased center: The unbiased center sets single-qubit angles to π/4 and two-qubit angles to zero, producing a uniform computational-basis distribution rather than a local maximum.Its correlators are centered at zero, so it generally lies in a region with non-negligible gradients unless target coefficients vanish or are exponentially small.
  • Identity center: At the identity center, gradient components vanish and the Hessian is diagonal with negative entries, making the point a local maximum driven by mismatch.The model-sensitivity term disappears because correlator derivatives vanish at the extremal correlator configuration.
  • Comparison: Identity and unbiased initialization can both have large local curvature for different reasons: mismatch at identity and maximal sensitivity of selected low-weight features at the unbiased center.The curvature decomposition separates target-model mismatch from squared model sensitivity.
  • Identity center: Identity initialization can provide inverse-polynomial curvature when at least one single-qubit target marginal remains system-size independent.For a low-body MMD loss with σ ∈ Θ(√n), one such marginal gives local curvature scaling as Θ(1/n) and a non-vanishing patch-variance guarantee.
  • Data-dependent center: Data-dependent marginal matching retains a nonzero sensitivity contribution after matching low-order target statistics, yielding local trainability under approximately factorizable targets.The guarantee depends on assumptions controlling residual mismatch while preserving model sensitivity; the authors note these assumptions may be stronger than necessary.

Appendix I: Numerical experiment for global MMD

The global-MMD experiment compares initialization schemes across scales and qubit counts. Despite global-MMD’s expected weak local signal, identity initialization shows the strongest favorable gradient scaling in the reported restricted patches.

  • Experimental design: The reported numerical comparison uses constant-bandwidth global MMD and includes both analytically studied schemes and covariance-based initialization for genomic data.The covariance scheme correlates two-qubit parameters and rescales perturbations using target-distribution covariances.
  • Global-MMD setting: Global MMD is expected to be difficult to train because local signal-bearing terms contribute only a small fraction of the uniformly distributed loss.For constant bandwidth, the weights remain broadly distributed across correlators, reducing the relative contribution of local terms.
  • Gradient behavior: Identity initialization yields surprisingly large gradients in restricted patches, while unbiased and data-dependent initialization show lower magnitude and less favorable scaling with qubit number.The authors interpret the identity result through its large curvature when the target differs from the all-correlator-one configuration.
  • Experimental design: Figure 7 varies initialization scale for identity, unbiased, data-dependent, and covariance schemes across circuits with n = 6 to n = 16 qubits.The figure also reports maximum variance over initialization scales and the scale achieving that maximum as functions of n.
Loading 2603.14576v1…