Source-linked AI summary
Probabilistic error cancellation with sparse Pauli-Lindblad models on noisy quantum processors
Ewout van den Berg, Zlatko K. Minev, Abhinav Kandala, Kristan Temme
TL;DR
Probabilistic error cancellation requires learning and inverting correlated noise, but representing and sampling the inverse of sparse Pauli models remains challenging. The paper develops a locally correlated, sparse Pauli-Lindblad approach with efficient noise learning and establishes conditions supporting its inversion.
Problem
Sampling the inverse of a sparse Pauli noise model requires accounting for exponentially many products of included Pauli terms.
Method
The paper learns locally correlated noise using sparse Pauli-Lindblad models containing single-qubit and selected neighboring two-qubit terms.
Results
The model-learning procedure can use nine measurement bases under mild qubit-topology conditions, and the associated matrix is full rank for fidelity-pair estimation.
Takeaways & Limitations
Sparse locally correlated models provide an efficient framework for representing noise and constructing its inverse for probabilistic error cancellation.
Abstract
from arXiv · showhide
Noise in pre-fault-tolerant quantum computers can result in biased estimates of physical observables. Accurate bias-free estimates can be obtained using probabilistic error cancellation (PEC), which is an error-mitigation technique that effectively inverts well-characterized noise channels. Learning correlated noise channels in large quantum circuits, however, has been a major challenge and has severely hampered experimental realizations. Our work presents a practical protocol for learning and inverting a sparse noise model that is able to capture correlated noise and scales to large quantum devices. These advances allow us to demonstrate PEC on a superconducting quantum processor with crosstalk errors, thereby providing an important milestone in opening the way to quantum computing with noise-free observables at larger circuit volumes.
Supplementary Information:
This section is identified by the paper title “Probabilistic error cancellation with sparse Pauli-Lindblad models on noisy quantum processors.”
- The paper is titled “Probabilistic error cancellation with sparse Pauli-Lindblad models on noisy quantum processors.”
- The title specifies sparse Pauli-Lindblad models as the paper’s modeling focus.
- The title places this work in the context of noisy quantum processors.
SUMMARY OF THE METHOD
The method uses processor topology to construct a sparse Pauli model, fits its parameters from benchmark-derived fidelities, and applies the learned quasi-probability distribution to mitigate circuit observables. For two-qubit layers, fidelity-product terms from two lists are combined in the model matrix.
- Processor qubits, gates, and topology define model Paulis containing all weight-one terms and selected weight-two terms on connected qubit pairs.
- Preparation for model fitting: Benchmark circuits at multiple depths are fit with exponential decays to estimate individual fidelities or fidelity-pair products, completed using unit-depth circuits or symmetry assumptions.
- Estimated fidelities form vector ˆf, while matrix M = M(B, K) is constructed and model parameters are set by the specified fitting problem.
- For a circuit containing the target layer, multiple instances sample preceding Paulis from the quasi-probability distribution, apply Pauli twirling, estimate observables, and scale them by γ.
- For two-qubit layers, the model replaces M with M(B1, K) + M(B2, K), with ˆf containing products of two fidelities.
BACKGROUND AND REVIEW … C. Scalable noise models
The paper reviews how noisy quantum-circuit operations are represented, simplified through Pauli twirling, and inverted quasi-probabilistically for unbiased observable estimation. It then introduces bounded-degree correlated Pauli channels as a route to reducing the exponential storage and processing burden of general models.
- BACKGROUND AND REVIEW: Quantum-circuit execution repeatedly initializes qubits, applies gates, and measures observables, motivating a representation of noisy operations as ideal operations combined with noise channels.The review frames each circuit execution as repeated preparation, gate application, and measurement, while noisy operations are modeled using an ideal operation and a noise channel.
- A. Noise channel simplification: Pauli twirling converts a generally dense noise channel into a diagonal Pauli-transfer representation that is more compact and directly invertible.The untwirled representation can require O(4^2n) coefficients, whereas twirling yields diagonal transfer-matrix entries corresponding to Pauli fidelities.
- A. Noise channel simplification: The diagonal fidelities are transformed into nonnegative coefficients defining a normalized Pauli-channel distribution, which can be implemented by sampling Pauli operators.The coefficients satisfy c_i ≥ 0 and sum to one; for Clifford gates, twirling is implemented by inserting sampled Paulis and their conjugates around the noisy gate.
- A. Noise channel simplification: Two-design twirling can simplify noise further to a single parameter because all nonidentity Pauli fidelities become equal.This is a broader alternative to Pauli twirling, with the Clifford group providing an example of a two-design.
- B. Quasi-probabilistic noise inversion: For a Pauli channel, probabilistic error cancellation targets the inverse transfer matrix, but its resulting coefficients are generally negative and therefore not a physical channel.The inverse has reciprocal fidelities on the diagonal; except when all fidelities equal one, the corresponding coefficients include negative values.
- B. Quasi-probabilistic noise inversion: Sampling Paulis from the absolute-value quasi-probability distribution, applying their signs and a normalization factor, produces an unbiased estimator of noiseless observables at increased variance.Each sampled circuit is evaluated, then its observable estimate is scaled by the sampled sign and γ; averaging these scaled values yields the ideal expectation value.
- C. Scalable noise models: Bounded-degree correlated Pauli channels factor their probability distributions using conditional independences, reducing the model complexity from storing and processing 4^n coefficients.This scalable model addresses the general exponential representation cost while retaining correlations of bounded degree [28].
PAULI-LINDBLAD NOISE MODEL … D. Variance in mitigated observable
The paper develops a valid sparse Pauli-Lindblad channel whose coefficients describe locally correlated noise, can be learned by nonnegative least squares, and can be inverted by coefficient negation. Error mitigation increases sampling requirements: maintaining fixed estimator variance requires trial counts proportional to γ^2.
- PAULI-LINDBLAD NOISE MODEL: The Lindbladian dynamics are represented as a conventional matrix exponential, and commuting Pauli terms allow the evolution operator to be factorized.The commutativity reflects that Pauli channels commute.
- PAULI-LINDBLAD NOISE MODEL: The model links measured Pauli fidelities to coefficients through a matrix M whose entries indicate whether Pauli operators commute or anticommute.The relation is expressed using the fidelity vector and the model-coefficient vector.
- PAULI-LINDBLAD NOISE MODEL: The Pauli-Lindblad model omits internal Hamiltonian dynamics, uses Pauli Lindblad operators, and remains a valid Pauli channel for all λ ≥0.Its coefficients are nonnegative, and the identity fidelity is always one.
- A. Channel operations: Successive Pauli channels combine by adding their coefficient vectors, while the inverse noise model is obtained by negating those coefficients.These properties follow from multiplicative Pauli fidelities and inverse fidelities.
- B. Sparse models: The model is sparse because gate noise is expected to have limited spatial range, making correlations beyond local neighborhoods negligible.The model uses a polynomial-size set of Pauli terms, |K| ≪ 4^n −1, selected to capture hardware correlations.
- C. Learning the model: The noise coefficients can be learned by fitting −log(f) with a nonnegative least-squares problem over the selected model Pauli terms.The fitting matrix has relatively few columns for a sparse model and generally far more rows than columns.
- D. Variance in mitigated observable: Keeping the variance of the mitigated-observable estimator fixed requires scaling the number of trials n proportional to γ^2.The estimator is formed from ±1 samples and rescaled by the mitigation factor γ.
NOISE LEARNING FOR SINGLE-QUBIT GATES WITH CROSSTALK
The protocol benchmarks single-qubit-gate layers by Pauli-twirling their noise, extracting individual Pauli fidelities, and fitting a two-local Lindblad model that captures crosstalk. Under mild topology conditions, the required fidelity measurements can be performed in nine bases, while dividing estimates across cycle counts removes state-preparation and readout errors.
- Noise learning for single-qubit gates with crosstalk: Dividing estimates obtained with k and zero cycles yields an unbiased estimate of f_i^k free of state-preparation and readout errors.Random Pauli conjugations implement a Pauli twirl, producing a Pauli channel with diagonal Pauli-transfer matrix and individual fidelities f_i.
- Noise learning for single-qubit gates with crosstalk: A two-local Lindblad model uses all weight-one Paulis and connected-support weight-two Paulis to capture crosstalk from single-qubit-gate layers.The model is fit from estimates of individual Pauli fidelities, selecting the Pauli set B subject to sampling-error limitations.
- Noise learning for single-qubit gates with crosstalk: Nine measurement bases suffice to estimate the required Pauli fidelities under mild conditions on the qubit topology.The model’s sample complexity and inverse accuracy are analyzed with fidelity estimates subject to sampling error that decreases as circuit instances increase.
A. Fidelities for model fitting
For two-local Lindbladian noise, model coefficients are assigned to Pauli supports on connected qubits and their subsets, including individual qubits. The resulting Pauli-basis matrix is full rank, with invertibility established through structured block elimination.
- Model construction: Two-local model coefficients include Paulis supported on connected qubits and all subsets of those supports, including individual-qubit Paulis.Qubit topology is represented as an undirected graph whose edges denote physical or logical connections.
- Fidelity identifiability: The Pauli-basis matrix M(P, P) is full rank for supports formed by taking every non-empty subset of the specified supports.The theorem defines P from all Pauli strings supported on the union of these subsets.
- Proof of invertibility: Invertibility is proved by ordering support sets by size and applying row and column sweeps that reduce the block matrix to invertible diagonal blocks.The sweep operations are applied when one support is contained in another, exploiting the block structure induced by support intersections.
B. Sample complexity and error analysis … NOISE LEARNING FOR TWO-QUBIT CLIFFORD GATES WITH CROSSTALK
The analysis bounds noise-model estimation from sparse fidelity measurements, reduces measurement overhead through nine Pauli bases, and characterizes complexity for selecting circuit depths. For two-qubit Clifford gates with crosstalk, twirling enables Pauli-channel learning, but some fidelities remain degenerate and require assumptions or alternative estimation methods.
- B. Sample complexity and error analysis: Sparse measurements of |B| ≪ 4^n − 1 fidelities are used to fit the full Pauli-Lindblad model and estimate its parameters under full-column-rank conditions.The analysis bounds deviations between estimated and ground-truth fidelities and parameters when measured fidelities satisfy the stated accuracy assumptions.
- 1. Measurement bases: Nine {X, Y, Z}⊗2 measurement bases suffice to cover all two-qubit Pauli bases on every edge when the qubit topology admits an ordering with at most two connected predecessors per vertex.This condition applies to commonly used two-dimensional grid and heavy-hexagon topologies.
- 2. Overall noise-learning complexity: Using nine bases at depths zero and k requires 18⌈N⌉ circuit instances, while unknown k can be selected by binary search over kmax in at most ⌈log2(kmax)⌉ trials.The failure probability is controlled by choosing δ = δ′/(|B| · ⌈log2(kmax)⌉).
- NOISE LEARNING FOR TWO-QUBIT CLIFFORD GATES WITH CROSSTALK: For arbitrary Clifford layers such as CX and CZ, Pauli twirling converts the associated noise into a Pauli channel whose fidelities can be probed through conjugation and engineered Pauli transformations.Invariant Paulis yield powers of individual fidelities, while non-invariant terms can sometimes be mapped back using additional single-qubit gates when support is preserved.
- NOISE LEARNING FOR TWO-QUBIT CLIFFORD GATES WITH CROSSTALK: Some Pauli transfers are degenerate, so cross terms reveal only products such as fIXfZX rather than the individual fidelities.The symmetry assumption extracts paired fidelities by treating them as equal, while assuming shared CZ and CX noise can infer additional fidelities from learned ones.
- NOISE LEARNING FOR TWO-QUBIT CLIFFORD GATES WITH CROSSTALK: Estimating individual fidelities by applying the noisy gate once is limited because initial and final Pauli components generally differ, preventing complete SPAM correction except for an initial ground state |0⟩.With paired fidelities, the learning equations instead use the sum M1 + M2 and the elementwise product f1 · f2.
A. Full rankedness of M when dealing with fidelity pairs
For fidelity-pair constructions under the stated non-overlapping two-qubit gate conditions, the combined matrix M = M(B1, K) + M(B2, K) is full rank. The proof establishes this by adding Pauli blocks and eliminating cross-block terms until the remaining blocks are full rank.
- Fidelity-pair construction: The construction uses weight-two P1 terms and P2 terms whose weights are one or two for gated pairs and up to four for ungated pairs.The model terms are organized into lists B1 and B2 containing fidelity-product pairs at corresponding locations.
- Full-rank theorem: Theorem SV.1 states that M = M(B1, K) + M(B2, K) is full rank.
- Proof for gated pairs: The proof begins with unit-weight Pauli blocks, where M(V, V) = I_n ⊗ Q is full rank, then adds gated-pair blocks without losing full rank.For a gated pair, elimination produces a lower-right block proportional to −2Q ⊗ Q, which is full rank.
- Proof for gated pairs: Combining matrices for each gated pair yields a diagonal scaling of Q ⊗ Q, so the sum remains full rank even when an individual matrix is not.The resulting diagonal entries are −2 and −4, and non-overlapping gates allow the procedure to be repeated independently.
- Proof for ungated pairs: For ungated pairs, row sweeps using single-qubit or incident-gate blocks reduce the combined lower-right block to −4Q ⊗ Q.The factorization of each Pauli term and the non-overlap condition determine which rows perform the eliminations.
PROBABILISTIC ERROR CANCELLATION AND ERROR-ANALYSIS · A. Sampling from the inverse
The protocol estimates observables by learning twirled Pauli-Lindbladian noise, implementing its inverse, and sampling that inverse quasi-probabilistically. Exploiting the sparse model’s product structure enables layerwise Pauli sampling while exposing overhead–complexity trade-offs.
- PROBABILISTIC ERROR CANCELLATION AND ERROR-ANALYSIS: PEC targets accurate expectation values of observables when each ideal circuit operation is accessible only through a noisy implementation.The observable is assumed to have operator norm ∥A∥≤1.
- PROBABILISTIC ERROR CANCELLATION AND ERROR-ANALYSIS: The method learns each twirled Pauli-Lindbladian channel as an estimate ˆΛi and implements its inverse ˆΛi^-1 experimentally.The inverse implementation follows the protocol described in section SVI A.
- A. Sampling from the inverse: Because the learned sparse Pauli model is not in canonical form, PEC samples its inverse by exploiting the product structure of commuting individual Pauli channels.For each included k, the channel factor has weight wk = (1 + e^-2λk)/2, and the overall inverse is the product of the individual inverses.
- A. Sampling from the inverse: For every k in K, sampling chooses identity with probability wk or Pk with probability 1 − wk, while recording a minus sign whenever Pk is chosen.The sampled Pauli terms are multiplied using their Abelian structure to form one full-inverse sample.
- A. Sampling from the inverse: Each sampled circuit output is multiplied by the accumulated sign and normalization factor γ, with the procedure applied at every circuit layer.The layer factors compound into the full sampling overhead γ(l) = Ql_i=1 γi.
- A. Sampling from the inverse: The mitigation procedure preserves the random-circuit instances, changing only the classical sampling distribution and postprocessing of measurement outcomes.A global sign flip is assigned from the total number m of sampled Pauli matrices as (−1)m.
- A. Sampling from the inverse: Expanding subsets of sparse-model terms can reduce γ, trading the computational cost of channel expansion against the resulting sample complexity.Combining terms permits Pauli channels with more terms and lowers the sampling overhead parameter.
B. Error bounds for probabilistic error cancellation · C. Weak exponential scaling
The error analysis bounds PEC estimation error by sampling noise and inaccuracies in learned noise channels, while a layer of k non-overlapping depolarizing gates yields weak exponential scaling through 15k sparse model coefficients.
- B. Error bounds for probabilistic error cancellation: The bound uses the estimated scaling factor γ(l), and the proof controls composed-channel discrepancies with diamond-norm submultiplicativity.The noise channel and its estimate are distinguished explicitly in the full error analysis.
- B. Error bounds for probabilistic error cancellation: The PEC error has two contributions: increased sampling error from quasi-probability mitigation and error from the noise-learning procedure.This decomposition motivates separating sampling accuracy from the discrepancy between ideal and estimated noise inversions.
- B. Error bounds for probabilistic error cancellation: The expectation-value error is bounded by γ(l)ϵs + ∥T_l − S_l∥⋄ for observables with ∥A∥≤1.The first term is the amplified sampling error, while the diamond-norm term captures the learned-channel mismatch.
- C. Weak exponential scaling: For k non-overlapping two-qubit depolarizing gates, the layer-level sparse noise model has 15k nonzero coefficients, each equal to −log(f)/16.The construction combines independent two-local error models at the layer level, with f denoting the Pauli fidelity.
- C. Weak exponential scaling: The resulting layer noise model therefore exhibits weak exponential scaling of the PEC overhead through the number of nonzero model coefficients.This conclusion follows by applying the scaling relation to the 15k-coefficient model for non-overlapping gates.
SETUP OF THE EXPERIMENT · A. Devices of the experiment
The experiments used fixed-frequency transmon qubits on heavy-hexagon superconducting processors, with main-text results obtained on ibm hanoi. Additional Falcon-chip tests produced similar results, supporting protocol reproducibility and robustness.
- A. Devices of the experiment: The processors employed fixed-frequency transmon qubits.
- A. Devices of the experiment: All devices were patterned with heavy-hexagon lattices.
- A. Devices of the experiment: Slight improvements in gate fidelity substantially reduce the circuit instances required for comparable estimation variance.
- A. Devices of the experiment: Main-text experiments used the 27-qubit Falcon processor ibm hanoi.
- A. Devices of the experiment: Other protocol iterations ran on ibm mumbai, ibm kolkata, ibm syndey, and ibm montreal.
- A. Devices of the experiment: These additional chips yielded results and conclusions similar to those obtained on ibm hanoi.
B. Specifications of the primary device · C. Dynamical decoupling
The primary device was IBM Hanoi, characterized by its gate implementation, coherence, readout performance, and temporal variability. Dynamical decoupling targeted idle-period decoherence and low-frequency noise, with a seven-qubit experiment examining its effect on learned noise structure.
- B. Specifications of the primary device: All circuits were transpiled to the standard basis gate set, and IBM Hanoi’s qubit topology and native CX durations were documented in Fig. S6.
- B. Specifications of the primary device: Single-qubit X gates used calibrated Gaussian DRAG pulses lasting 35.5 ns, while I and RZ gates were virtualized and CX gates used optimized cross-resonance pulses.
- B. Specifications of the primary device: IBM Hanoi had quantum volume 64, with average energy-relaxation T1 and Hahn-echo T2 times of 151 µs and 107 µs, respectively.
- B. Specifications of the primary device: The device’s coherence times fluctuated over two months across all 27 qubits, so mitigation experiments were interleaved with noise-learning runs every few hours.
- B. Specifications of the primary device: The average assignment readout error across all device qubits was 2.5%, with energy relaxation biasing P(1 | 0) below P(0 | 1).
- C. Dynamical decoupling: Different CX gate times created idle periods, especially for context qubits, so dynamical decoupling applied a standard Xp −Xm sequence to reduce decoherence and low-frequency noise.
- C. Dynamical decoupling: A seven-qubit layer with two CX gates and three idle context qubits was learned with and without dynamical decoupling to assess changes in noise structure.
- C. Dynamical decoupling: Without dynamical decoupling, the dominant learned noise consisted of unit-weight Pauli-Z terms.
D. Additional learning and control experiments
Additional experiments provide the complete nine-basis learning data and show that unit-depth and symmetry-based noise-model fitting agree overall, despite localized differences. They also examine learned coefficients with and without dynamical decoupling.
- Additional learning experiments: Unit-depth and symmetry-based fitting produce noise-model profiles that match well overall, with only localized differences for a 20-qubit layer containing 10 cx gates.The comparison uses post-processing methods applied to the same layer; symmetry-based fitting was used for the other experiments in the work.
- Control experiments: Dynamical-decoupling experiments compare learned noise-model coefficients with and without an Xp −Xm sequence applied to idle context qubits during a concurrent-cx layer.The sequence is inserted during idle times, and the figures identify the supports of the resulting model coefficients.
- Additional learning experiments: The full learning data for the Fig. 2a setup spans all nine bases prescribed by the protocol.Fig. S10 reports observable expectations across circuit depths with exponential-decay fits for the four-qubit layer on ibm hanoi.