Source-linked AI summary

A Theory of Finite-Noise Optima and Generalization in Quantum Machine Learning

Ziyu Zhang, Zikang Jia, Xiaosong Li, Yulong Dong

arXiv:2608.24229v1quant-phcs.LG

TL;DR

Quantum noise’s intermediate effect on quantum machine-learning performance is not explained by weak-noise error accumulation or strong-noise trainability collapse. This paper develops a statistical learning theory showing that noise-order purity links noise to complexity and generalization, producing a finite-noise optimum.

  • Problem

    The intermediate-noise regime between perturbative error accumulation and strong-noise trainability collapse remains insufficiently explained.

  • Method

    A noise-order surrogate introduces purity as a measure of response mixing, linking noise-induced complexity reduction and generalization-gap reduction against increasing prediction bias.

  • Results

    The analysis and numerical experiments show that competing complexity and bias effects produce a setup-dependent finite-noise optimum and explain nonmonotonic testing performance.

  • Takeaways & Limitations

    Noise programming can move an under-regularized noisy model toward its finite-noise optimum.

  • Takeaways & Limitations

    The simulations assume independent gate-level noise, leaving correlated or time-dependent errors and finite-shot output estimation for future validation.

Abstract

from arXiv · show

Quantum noise is expected to degrade quantum machine learning by driving circuits away from their noiseless implementations. Yet recent studies show moderate noise can reduce testing error, a behavior unexplained by weak-noise perturbative error accumulation or strong-noise trainability collapse. Here we develop a statistical learning theory connecting microscopic noise processes to macroscopic learning performance. At its heart is a noise-order purity parameter, derived from a surrogate model analysis, that predicts the noise-induced reduction in model complexity and the consequent reduction in the generalization gap. Noise simultaneously increases prediction bias. Their competition explains the intermediate-noise regime left open between these limits. It produces a finite-noise optimum whose location depends on the learning setup and can disappear in the large-sample limit. Numerical experiments validate these predictions. Noise programming can move a model towards this optimum. These results make the non-monotonic effect of noise predictable and provide a route to harness it.

I. INTRODUCTION … II.2. A noise-order surrogate explains noise-induced regularization

The paper develops a statistical learning theory explaining why finite quantum noise can improve generalization despite worsening training fit. A noise-order surrogate attributes this effect to reduced sensitivity across noise orders, while prediction distortion eventually degrades performance.

  • I. INTRODUCTION: Quantum noise need not monotonically harm learning because a predictor adapts to trainable components, finite data, and physical noise rather than reproducing a prescribed noiseless computation.Noise can remove useful signal while also suppressing excess freedom used to fit the training sample.
  • I. INTRODUCTION: The paper addresses the unresolved question of when noise improves testing, what determines the optimal noise level, and why the benefit eventually disappears.Existing weak-noise bounds control worst-case deviations, whereas strong-noise analyses focus on representation and trainability loss.
  • I. INTRODUCTION: The proposed theory organizes noisy responses by noise occurrences and uses noise-order purity to connect noise-induced model-complexity reduction with the generalization gap.Noise also attenuates the learnable signal, creating a competition between regularization and prediction bias.
  • II.1. Finite noise induces non-monotonic generalization: In the diabetes regression experiment, training loss generally increases with physical noise across all four channels, while test loss first falls below the nearly noiseless value and later rises.The experiment fixes the dataset split, four-qubit architecture, and training protocol while varying the physical noise rate β.
  • II.1. Finite noise induces non-monotonic generalization: The intermediate-noise improvement reflects changed generalization rather than better optimization or a lower training loss.The observed regime occurs after noise changes generalization but before strong noise suppresses useful gradients and destroys trainability.
  • II.1. Finite noise induces non-monotonic generalization: The normalized generalization gap decreases rapidly in the same intermediate-noise regime where test loss improves, before large-noise degradation begins.This establishes finite quantum noise as a mechanism that can reduce the discrepancy between training and test performance.
  • II.2. A noise-order surrogate explains noise-induced regularization: The surrogate groups microscopic noise realizations by total noise order, with binomial weights distributing probability across effective predictors as β increases.Near β=0, probability concentrates on the nearly noiseless branch; larger β spreads weight across more finite-noise branches.
  • II.2. A noise-order surrogate explains noise-induced regularization: Its loss separates distortion of the averaged noisy predictor from a stochastic penalty on disagreement across noise orders, explaining worse training fit alongside improved intermediate-noise generalization.The first term reflects hypothesis-space distortion, while the second discourages strong sensitivity to realized noise order.

II.3. Noise-order purity predicts effective complexity

Noise-order purity measures how probability spreads across noise-order predictors: it decreases as finite noise broadens the ensemble, producing stronger averaging and lower effective model complexity. Through this complexity reduction, purity predicts the noise dependence of the generalization gap while noise also induces regularization against predictor disagreement.

  • Noise-order purity predicts effective complexity: The surrogate noisy QML model is a weighted ensemble whose binomial weights depend only on physical noise rate β and noisy locations g.As β increases, probability mass spreads from the nearly noiseless branch across finite-noise branches.
  • Noise-order purity predicts effective complexity: Noise-order purity is near one when one noise order dominates and decreases as probability mass spreads across many finite-noise orders.Lower purity therefore indicates stronger averaging across noisy circuit responses.
  • Noise-order purity predicts effective complexity: The purity decays across regimes, following s2(β, g) = 1−2gβ +O(g2β2) for gβ ≪1 and s2(β, g) ≈(4πgβ)−1/2 when gβ ≫1 and β ≤1/2.Its regime-dependent shape reflects the transition from the nearly noiseless branch to a broad finite-noise distribution.
  • Noise-order purity predicts effective complexity: When noise-order predictors are weakly correlated, stochastic noise regularizes disagreement among them, and effective model complexity is proportional to noise-order purity.The proportionality is derived under idealized assumptions but validated numerically across the studied QML models.
  • Noise-order purity predicts effective complexity: As purity decreases, effective fitted degrees of freedom fall, so the generalization gap’s β-dependence is controlled by noise-order purity when data-noise variance and ntrain remain fixed.This relationship agrees with the observed decay of ∆gen(β) in Fig. 1(e).

II.4. Effective-complexity scaling in trained models … II.7. Noise-induced regularization persists in molecular QML

The paper links physical noise to reduced effective complexity and generalization-gap reduction, while stronger noise causes trainability collapse. Controlled parameter perturbations can emulate increased effective noise, and the same finite-noise regularization mechanism persists in molecular QML.

  • II.4. Effective-complexity scaling in trained models: The effective-complexity reduction is directly tested using the Jacobian of the trained noisy predictor and compared with noise-order purity.The decreasing generalization gap provides indirect evidence for the predicted complexity reduction.
  • II.4. Effective-complexity scaling in trained models: All four noise models reduce effective model complexity at different rates, and normalized complexity follows the predicted dependence on s2(β, geff).The effective-noise parameter geff is fitted separately for each noise model.
  • II.5. Finite-noise improvement precedes trainability collapse: For unital noise models, output-gradient norms decrease relatively uniformly, whereas amplitude damping suppresses parameters more strongly with greater distance from the final measurement.The distance dependence is consistent with results on quantum machine learning under non-unital noise.
  • II.5. Finite-noise improvement precedes trainability collapse: In the intermediate regime II, noise reduces effective complexity while output-gradient norms remain appreciable, distinguishing regularization from trainability collapse.Stronger noise in regime III suppresses trainability, increasing both training and test errors.
  • II.6. Noise programming through controlled parameter perturbations: Noise programming adds independent Gaussian perturbations with standard deviation σ to parameters after each optimization step, modeling uncertainty in variational gate parameters.The method is introduced because a device’s native noise level is largely fixed.
  • II.6. Noise programming through controlled parameter perturbations: Programming can mimic increased effective noise at fixed physical noise, but stronger programming gradually erases the non-monotonic response through high effective noise.Figure 4 maps testing error over physical noise rate β and programming strength σ, with “none” as the unprogrammed baseline.
  • II.7. Noise-induced regularization persists in molecular QML: Across molecular QML comparisons, testing MSE consistently has a finite-noise optimum, while training MSE increases with noise rate, indicating increasing noise-induced bias.The comparisons span different training-set sizes, noise models, and molecular datasets.
  • II.7. Noise-induced regularization persists in molecular QML: Molecular QML results confirm linear generalization-gap scaling with effective complexity and proportionality between effective complexity and noise-order purity.Finite noise reduces complexity and can improve testing performance before training error becomes dominant.

II.8. Finite-noise optima disappear in the large-sample limit … IV.1. QML models and datasets

Finite-noise benefits are conditional: increasing training data shifts the optimum toward zero and can eliminate nonmonotonic testing responses. The discussion frames noise-order purity as predictive while emphasizing limits, open questions, and noise programming as a possible control strategy.

  • II.8. Finite-noise optima disappear in the large-sample limit: For ntrain = 80, 160, 320, and 1000, the three smaller sets retain finite-noise testing-error reductions, whereas ntrain = 1000 almost eliminates the nonmonotonic response.The ntrain = 1000 case also has a much smaller low-noise generalization gap.
  • II.8. Finite-noise optima disappear in the large-sample limit: As ntrain increases, the finite-noise optimum generally shifts toward the low-noise limit; once ntrain exceeds q, it moves toward zero.The passage attributes this shift to removal of the data limitation.
  • II.8. Finite-noise optima disappear in the large-sample limit: In the infinite-training-set limit, the generalization gap vanishes because the model sees the full data distribution and the test set adds no statistical information.Increasing training error with ntrain accompanies the shrinking generalization gap in Fig. 6(a).
  • III. DISCUSSION: The theory connects microscopic noise processes to macroscopic learning performance through noise-order purity, explaining nonmonotonic testing performance beyond weak- and strong-noise limits.Noise-order purity measures how broadly the model response is mixed across noise orders.
  • III. DISCUSSION: Nonmonotonic noise effects are presented as general and predictable rather than occasional artifacts.The discussion states that noise is not intrinsically beneficial and that improvements are not unconditional.
  • III. DISCUSSION: The finite-noise optimum differs between over- and underparameterized models and can disappear in the large-sample limit.The analysis combines a noise-order surrogate with local linearization around the trained optimum.
  • III. DISCUSSION: Noise programming can move an under-regularized noisy model toward its finite-noise optimum.The discussion proposes allocating protection according to how circuit components contribute to predictive information.
  • III. DISCUSSION: The simulations assume independent gate-level noise, leaving correlated errors, time dependence, and finite-shot output estimation as hardware questions for future work.The open question is whether noise-order purity remains predictive under these more realistic conditions.

IV.1.1. Four-qubit regression model · IV.1.2. Molecular QGNN models · IV.2. Quantum noise models

The experiments use a four-qubit diabetes regression circuit and molecular QGNNs for fixed-size molecular regression, evaluating four single-qubit noise channels applied throughout quantum circuits. The molecular models use eight or nine qubits with trainable classical readouts, while datasets and simulation procedures are fixed across key experimental conditions.

  • IV.1.1. Four-qubit regression model: The diabetes regression experiment uses body mass index and log serum triglyceride level from 442 samples, with disjoint training and testing sets of 40 and 400 samples.Both input features are independently rescaled to [0, π], and the regression target is rescaled to [−1, 1] using training-set-fitted transformations.
  • IV.1.1. Four-qubit regression model: The four-qubit circuit has nine layers, re-uploads both features through RX rotations, applies periodic-ring IsingXX interactions, and uses one RY rotation per qubit.Its scalar output is the expectation value of Z⊗4 under the circuit dynamics.
  • IV.1.2. Molecular QGNN models: Molecular regression targets the HOMO–LUMO gap in electronvolts using QM9-HA8 molecules with eight heavy atoms and PCQM4Mv2-HA9 molecules with nine heavy atoms.The datasets are studied on fixed-size molecular subsets.
  • IV.1.2. Molecular QGNN models: The molecular experiments use fixed ordered training pools of 2,000 and 320 molecules and disjoint testing sets of 1,500 molecules, with nested training prefixes and fixed tests.These testing sets remain fixed across training sizes, noise rates, and runs.
  • IV.1.2. Molecular QGNN models: Each heavy atom maps to one qubit, producing eight- and nine-qubit QGNNs with rotation-based atom encoding, bond-dependent two-qubit interactions, and a trainable linear readout.The circuit uses one EDU-QGC-inspired graph-circuit layer and extracts the complete computational-basis probability vector.
  • IV.1.2. Molecular QGNN models: The QM9-HA8 and PCQM4Mv2-HA9 models contain 310 and 569 trainable parameters in total, respectively.The component counts are 53 quantum-circuit plus 257 classical readout parameters for QM9-HA8, and 56 plus 513 for PCQM4Mv2-HA9.
  • IV.2. Quantum noise models: The study considers depolarizing, bit-flip, phase-damping, and amplitude-damping single-qubit noise channels with physical noise rate β ∈[0, 1].Each selected channel is applied after every elementary gate in data encoding and trainable circuits, independently to both qubits after two-qubit gates.
  • IV.2. Quantum noise models: Noise is excluded from classical post-processing, while noiseless circuits use state-vector simulation and noisy circuits use PennyLane density-matrix simulation.The four-qubit circuit has a fixed number of no…

IV.3. Training and evaluation · IV.4. Local model complexity and effective noise count

Training uses fixed optimization and evaluation protocols with physical noise present during both phases, while testing data remain held out. Local complexity is estimated from predictor-Jacobian eigenvalues using a soft effective count and normalized noise-order fits.

  • IV.3. Training and evaluation: The four-qubit model minimizes mean-squared error with full-batch Adam at learning rate 0.01 for 800 epochs without early stopping.The selected physical noise channel is included during training and evaluation, while testing samples do not update parameters.
  • IV.3. Training and evaluation: Four-qubit experiments train ten independently initialized models, reusing the same ten initial parameter vectors across every noise rate and model.Each run samples 72 initial parameters from a standard normal distribution using a fixed random seed, and metrics are evaluated after the final update.
  • IV.3. Training and evaluation: Molecular models minimize mean-squared error in electronvolts squared using Adam with learning rate 0.03, mini-batches of 16, and gradient clipping at maximum norm 1.QM9-HA8 and PCQM4Mv2-HA9 are trained for 300 and 100 epochs, respectively, with metrics evaluated at the final checkpoint.
  • IV.3. Training and evaluation: Molecular noise sweeps reuse five fixed initializations, include physical noise during training and evaluation, and evaluate every final checkpoint on all 1,500 testing molecules.Curves report means over five runs with uncertainty shown as one sample standard deviation.
  • IV.4. Local model complexity and effective noise count: Local model complexity is estimated by linearizing the physical noisy predictor around trained parameters, making the predictor Jacobian the design matrix of a local linear regression.For molecular models, the Jacobian covers the complete parameter vector, including quantum-circuit and linear-readout parameters.
  • IV.4. Local model complexity and effective noise count: Complexity uses a soft eigenvalue count rather than a hard nonzero count, because eigenvalues quantify response strength and numerical thresholds can distort hard counting.Directions with eigenvalues well above 10^-4 contribute approximately one, while weaker directions contribute proportionally less on a common scale.
  • IV.4. Local model complexity and effective noise count: Effective noise count is obtained by normalizing each run by its own zero-noise complexity, averaging, and fitting the result to exact binomial noise-order purity.The fit uses R = 10 for four-qubit experiments and R = 5 for molecular experiments; Fig. 1(e) instead uses exact g = 126 as its dashed reference.
  • IV.4. Local model complexity and effective noise count: The Fig. 3 parameter-wise output-gradient norm is computed from the same predictor Jacobian, with scalar-response profiles averaged over parameters and runs and normalized to β = 0.The supplied method description also specifies plotting parameter profiles as P_i(β)/10 and dividing scalar-response averages by their corresponding zero-noise average.

IV.5. Noise programming · DATA AVAILABILITY

Noise programming adds independently sampled Gaussian perturbations after each Adam update, while preserving other training conditions and evaluating final testing MSE across repeated retraining runs. The study also makes datasets, processed data, splits, and trained checkpoints publicly available.

  • IV.5. Noise programming: Noise programming adds an independent Gaussian perturbation after every Adam update.The perturbation is sampled independently for every update and every run.
  • IV.5. Noise programming: The programmed parameters follow an update rule defined using the Adam update ΔAdam_t and the sampled perturbation.The supplied passage identifies ΔAdam_t as the optimizer update at step t, while the remainder of the equation is not included.
  • IV.5. Noise programming: Programming strengths are σ = 2−k for k = 1, . . . , 10, alongside an unprogrammed baseline σ = 0.Initial parameters, physical noise channel, and all other training settings remain fixed.
  • IV.5. Noise programming: For every physical noise rate β and programming level k, models are retrained for 800 full-batch epochs in each of ten runs.Each programming-landscape cell reports the final testing MSE averaged across those runs.
  • DATA AVAILABILITY: The diabetes dataset is distributed with scikit-learn, while QM9 and PCQM4Mv2 are publicly available from their original sources.The passage identifies the cited sources as [43],, and [63], respectively.
  • DATA AVAILABILITY: Fixed molecular subsets, data splits, processed numerical data underlying the figures, and trained checkpoints are available in the study’s GitHub repository.Repository: https://github.com/dongsnaq/Finite-Noise-Generalization-QML.

Appendix A: Theory of noise-induced regularization · A.1. Problem setup and stochastic noise-occurrence expansion

This appendix models noisy QML circuits through stochastic noise occurrences and expands the noisy predictor by occurrence count. The resulting surrogate separates mean-channel error from realization disagreement, showing how noise filters the accessible function class and induces regularization.

  • A.1. Problem setup and stochastic noise-occurrence expansion: The setup represents a QML model as parametrized gates acting on an input density matrix and followed by observable measurement.Data are assumed to enter through the initial state, while more general data-dependent gate maps only modify model-dependent quantities.
  • A.1. Problem setup and stochastic noise-occurrence expansion: Noise is modeled as an additional quantum channel inserted after each ideal gate, producing a noisy output and predictor.The stochastic noise-occurrence construction is introduced to explain non-monotonic noise response beyond the simplest example.
  • A.1. Problem setup and stochastic noise-occurrence expansion: In the simplest model, each location either leaves the circuit unchanged or applies an additional noisy operation, generating effective circuits with one or more noise occurrences.This includes random Pauli operations such as dephasing or bit-flip noise.
  • A.1. Problem setup and stochastic noise-occurrence expansion: Bernoulli variables z_j ∼ Ber(β) determine whether noise is inserted at each gate location, with the resulting conditional predictors averaged over noise realizations.The occurrence count m groups realizations according to how many noisy operations they contain.
  • A.1. Problem setup and stochastic noise-occurrence expansion: The noisy mean predictor expands over occurrence count as a binomially weighted combination of occurrence-m predictors.Effective predictors generated by different noise counts form the expansion basis, while the binomial occurrence distribution supplies the weights.
  • A.1. Problem setup and stochastic noise-occurrence expansion: The stochastic surrogate loss exposes an explicit variance term across noise realizations through the bias-variance identity.Its first term is the mean-channel loss of the averaged noisy predictor, and its second term penalizes disagreement among realization-specific predictors.
  • A.1. Problem setup and stochastic noise-occurrence expansion: The ensemble average ϕ_β = Σ_m p_mψ_m filters the accessible function class, providing a mechanism for noise-induced regularization.The same occurrence structure therefore affects both the noisy mean predictor and the surrogate variance penalty.

A.2. Noise-induced regularization in the surrogate loss · A.3. Noise-induced generalization effect and reduced effective model complexity

Noise-induced regularization penalizes disagreement among predictors indexed by noise occurrence counts, with its strength governed by the noise-order purity s2(β, g). In the surrogate model, lower purity reduces effective complexity and can narrow the generalization gap, while stronger noise increases training loss.

  • A.2. Noise-induced regularization in the surrogate loss: The stochastic surrogate loss includes a pairwise disagreement penalty among predictors associated with different noise occurrence counts.The covariance matrix Λ(β, g) is positive semidefinite.
  • A.2. Noise-induced regularization in the surrogate loss: Noise-order purity s2(β, g) measures concentration in the noise occurrence-count distribution.Its complement, 1 − s2(β, g), measures the distribution’s total variance.
  • A.2. Noise-induced regularization in the surrogate loss: When occurrence counts spread across more values of m, s2(β, g) decreases and the covariance matrix captures greater total variance.The same concentration quantity also controls the local sensitivity of the averaged noisy predictor.
  • A.3. Noise-induced generalization effect and reduced effective model complexity: The occurrence-count coarsened surrogate represents each noise realization with c(z) = m by a predictor ψm to expose noise-occurrence averaging’s effect on complexity.This is the surrogate model used for the generalization analysis.
  • A.3. Noise-induced generalization effect and reduced effective model complexity: For the four-qubit model, the parameter count is q = 72 while the number of noisy locations is g = 126.The analysis distinguishes trainable parameters from noisy locations; they need not coincide.
  • A.3. Noise-induced generalization effect and reduced effective model complexity: The local hat matrix H(θ; β, g) generally is not an exact projector, but its trace acts as the effective number of fitted degrees of freedom.The effective model complexity is defined as deff := Tr(H).
  • A.3. Noise-induced generalization effect and reduced effective model complexity: Under the simplified overlap model, deff = Tr(H) ≈ q (κ + (1 − κ)s2(β, g)), linking effective complexity to purity and predictor alignment.Here κ ∈ [0, 1] measures average alignment between different occurrence-count predictors.
  • A.3. Noise-induced generalization effect and reduced effective model complexity: As the occurrence-count distribution spreads, the stochastic surrogate’s effective complexity decreases by a factor s2(β, g), potentially reducing the generalization gap while stronger noise increases training loss.This establishes the competing complexity and loss effects of noise in the surrogate model.

A.4. Scaling regimes of the noise-order purity · A.5. Mean-squared prediction error in the presence of quantum noise

Noise-order purity predicts how quantum noise reduces effective complexity: it remains near one under weak noise but decreases in the spread-out regime, lowering complexity and the generalization gap. In the MSE, this benefit competes with noise-induced bias, producing a non-monotonic testing error and a setup-dependent optimum.

  • A.4. Scaling regimes of the noise-order purity: Noise-order purity s2(β, g) approximately reduces effective model complexity and equals one at zero noise.As noise strength increases, the distribution spreads over more noise orders and s2(β, g) decreases.
  • A.4. Scaling regimes of the noise-order purity: When gβ ≪1, the noise-order distribution concentrates near zero, so noise produces only a perturbative reduction in effective complexity.The weak-noise regime is characterized by gβ ≪1.
  • A.4. Scaling regimes of the noise-order purity: When gβ ≫1, the binomial noise-order distribution is approximated by a Gaussian with mean µ = gβ and variance σ2 = gβ(1 −β).The spread-out distribution yields a purity scaling proportional to 1/√(gβ(1 −β)).
  • A.4. Scaling regimes of the noise-order purity: The interpolating expression recovers s2(β, g) = 1 −2gβ + O(g2β2) for gβ ≪1 and s2(β, g) ≈(4πgβ)−1/2 for gβ ≫1 and β ≪1.A Poisson approximation with rate gβ also captures the transition between regimes, with I0 denoting the modified Bessel function of the first kind.
  • A.4. Scaling regimes of the noise-order purity: Noise can reduce effective complexity from O(g) to O(√g), reducing the stochastic surrogate’s generalization gap by O(1/√g).This reduction in effective complexity and generalization gap explains why moderate noise can improve testing performance.
  • A.5. Mean-squared prediction error in the presence of quantum noise: The training MSE contains an unlearnable component, a noise-induced bias that increases as purity s decreases, and a data-noise contribution; therefore, it increases with noise rate.The local regression surrogate acts as a local smoother under the weak-overlap approximation, with learnable dimension rank(PT ) = q.
  • A.5. Mean-squared prediction error in the presence of quantum noise: As noise increases, the generalization gap decreases while the bias term (1 −s)2 f∥ increases, so their competition produces a non-monotonic testing MSE.The testing MSE is the sum of training MSE and generalization gap; the gap cancels the training-set data-noise fitting term −2σ2qs/n.
  • A.5. Mean-squared prediction error in the presence of quantum noise: Noise is more likely to improve testing MSE when the learnable signal is weak, data noise is large, local model dimension is large, or the training set is small.The testing MSE minimum and corresponding optimal noise rate β⋆apply when the solution lies in the achievable range of s2(β, g).

A.6. Extensions beyond the Bernoulli noise model

The noise-order analysis extends beyond Bernoulli identity-versus-noisy operations to general principal channels and multi-branch mixtures, while preserving the core path-probability intuition.

  • General principal channels: For two-channel noise, each location samples E_j,z_j = (1 − z_j)A_j + z_jB_j with z_j ∼ Ber(β), and the mean channel is (1 − β)A_j + βB_j.A_j and B_j are quantum channels, with β controlling the Bernoulli branch selection.
  • General principal channels: Replacing the identity branch with a general principal channel preserves the noise-order expansion but distorts the model’s hypothesis space.The principal branch becomes a gate sequence distorted by the channel sequence A_j rather than the noiseless implementation.
  • Multi-branch noise: Multi-branch noise replaces binomial weights with path probabilities over branch-choice sequences across the circuit.A generalized noise-order purity is defined by summing the squared probabilities of these paths.

Appendix B: Numerical validation of the noise-order surrogate … B.3. Efficient evaluation of noisy order-resolved quantities by dynamic programming

Appendix B numerically validates the noise-order surrogate: surrogate complexity decreases with noise and follows the predicted noise-order purity scaling, while dynamic programming makes order-resolved evaluation polynomial rather than exponential in the number of noisy locations.

  • B.1. Surrogate effective complexity in trained models: The surrogate complexity decreases with noise for all four noise models and shares the same noise-order purity s2(β, geff) after normalization.This reproduces the s2 scaling predicted by Eq. (A33) and observed for the physical-model complexity.
  • B.1. Surrogate effective complexity in trained models: The surrogate explicitly penalizes disagreement among noise-order predictors, whereas the physical model captures the resulting regularization through the averaged noisy predictor.Consequently, the same physical noise produces stronger effective regularization in the surrogate and a larger fitted geff.
  • B.1. Surrogate effective complexity in trained models: A matrix order inequality shows that surrogate model complexity is no greater than physical-model complexity, confirming that the surrogate makes noise-order regularization explicit.The surrogate retains variation among order-resolved Jacobians in its local curvature.
  • B.2. Interpretation of the effective noise count: Phase damping has the smallest fitted geff because it often leaves measured populations unchanged and affects the predictor only when later gates convert lost coherence into population differences.Amplitude damping, bit-flip, and depolarizing noise can perturb the measured output more directly or broadly.
  • B.2. Interpretation of the effective noise count: The fitted geff summarizes each noise model’s accumulated effect on the measured predictor Jacobian, which depends on trained parameters, input states, and the measured observable.Remaining circuit gates can rotate earlier errors between population and coherence components.
  • B.3. Efficient evaluation of noisy order-resolved quantities by dynamic programming: Directly expanding noise channels across g locations requires 2^g event placements, whereas dynamic programming aggregates all placements of each order into one averaged state.Linearity makes averaging before subsequent propagation exactly equivalent to propagating every placement separately and averaging afterward, without assuming equivalent locations.
  • B.3. Efficient evaluation of noisy order-resolved quantities by dynamic programming: Retaining all noise orders costs O(nqg^2d^3) time, polynomial in g rather than the O(2^g) cost of enumerating placements.Order-resolved tangents follow the same recurrence as states, and negligible-binomial-probability sectors can be omitted.

Appendix C: Molecular QML simulation details and ablation studies … C.4. Ablation of noise models

The appendix details molecular QML circuits and ablations showing that optimization endpoint, classical post-processing, and noise channel influence finite-noise behavior. Across models, noise can reduce complexity and testing error, but channel-dependent rates and measurement basis shape the response.

  • C.1. Circuit structure and other simulation details: The molecular predictors use 310 or 569 trainable parameters, combining an eight- or nine-qubit quantum circuit with a linear probability-to-target map.The circuits contain 53 and 56 quantum parameters, while the classical maps add 257 and 513 parameters, respectively.
  • C.1. Circuit structure and other simulation details: For molecular graphs, noisy locations arise from atom, node, and bond operations, while geff summarizes noise-driven complexity reduction across the dataset.The physical channel acts at every executed location, and the location count depends on the molecular graph G.
  • C.2. Ablation of optimization epochs: Longer optimization lowers training error but increases low-noise testing error and generalization gap, while the overall non-monotonic response remains stable through 300 epochs.The comparison evaluates the same five QM9-HA8 trajectories after 100, 200, and 300 epochs.
  • C.3. Ablation of more complex classical post-processing: Replacing the linear readout with 128 tanh hidden units raises the total parameter count from 310 to 33,078 without improving the minimum testing error.Both readouts fit low-noise training data closely and retain lower testing error at finite noise, although the nonlinear response is more variable.
  • C.3. Ablation of more complex classical post-processing: Despite their different classical parameter counts, linear and nonlinear readouts have comparable effective full-model and quantum-only complexities across the noise sweep.The results indicate that noisy quantum representations primarily govern complexity reduction rather than added classical capacity.
  • C.4. Ablation of noise models: Bit flip, depolarizing, and amplitude damping produce intermediate-noise testing-error reductions, whereas phase damping does not show the same non-monotonic response.For the first three channels, testing minima occur before strong noise prevents accurate training-data fitting.
  • C.4. Ablation of noise models: All four channels reduce effective model complexity at different rates, so identical β values produce different complexity reductions and minima at different physical noise rates.The fitted geff maps physical noise to the noise-order purity s2(β, geff) for normalized complexity comparisons.
  • C.4. Ablation of noise models: Under phase damping, noise raises training error without monotonically reducing the generalization gap, and changing from Z- to X-basis readout tests whether measurement basis causes this behavior.The supplied passage states that the X-basis sweep recovers high-noise collapse and a linear relation, but its conclusion is truncated.
Loading 2608.24229v1…