Source-linked AI summary
Quantum Error Mitigation
Zhenyu Cai, Ryan Babbush, Simon C. Benjamin, Suguru Endo, William J. Huggins, Ying Li, Jarrod R. McClean, Thomas E. O'Brien
TL;DR
Quantum error mitigation is studied across diverse techniques, with attention to their implementation, performance, and limitations. The review surveys these methods and reports that their lower hardware requirements and broad applicability have made them integral to recent experimental demonstrations.
Problem
Comparing quantum error mitigation methods remains difficult because estimators involve different bias–variance trade-offs, resource requirements, and problem-dependent performance metrics.
Method
The review surveys quantum error mitigation concepts, implementation details, specific techniques, and comparisons, including zero-noise extrapolation and learning-based methods.
Results
At a fault rate of 10^-3 with sampling overhead 100, probabilistic error cancellation reduces physical-qubit overhead by 80% for some classically intractable problems and by more than 45% in many practical fault-tolerant applications.
Takeaways & Limitations
Quantum error mitigation has lower hardware requirements than quantum error correction and a broad range of techniques targeting diverse application scenarios.
Takeaways & Limitations
Standard expectation-value methods can incur sampling costs that grow exponentially with circuit fault rate and may lack the parallelisability of unmitigated estimation.
Abstract
from arXiv · showhide
For quantum computers to successfully solve real-world problems, it is necessary to tackle the challenge of noise: the errors which occur in elementary physical components due to unwanted or imperfect interactions. The theory of quantum fault tolerance can provide an answer in the long term, but in the coming era of `NISQ' machines we must seek to mitigate errors rather than completely remove them. This review surveys the diverse methods that have been proposed for quantum error mitigation, assesses their in-principle efficacy, and then describes the hardware demonstrations achieved to date. We identify the commonalities and limitations among the methods, noting how mitigation methods can be chosen according to the primary type of noise present, including algorithmic errors. Open problems in the field are identified and we discuss the prospects for realising mitigation-based devices that can deliver quantum advantage with an impact on science and business.
I. INTRODUCTION
Quantum error mitigation (QEM) targets noise-induced bias in near-term quantum computations by post-processing ensembles of noisy circuit runs. This review surveys QEM methods, their trade-offs and requirements, while emphasizing that mitigation remains practical only within hardware- and circuit-size boundaries.
- I. INTRODUCTION: Fault-tolerant quantum computation is theoretically possible below an error threshold, but current fault-tolerant implementations impose daunting qubit overheads.Examples cited include hundreds of thousands of qubits for some classically intractable scientific applications and millions for industrial applications.
- I. INTRODUCTION: QEM aims to convert continuing hardware improvements into immediate gains in quantum information processing before fully fault-tolerant systems are available.The review presents mitigation as practically effective while accepting that hardware imperfections limit algorithmic complexity.
- I. INTRODUCTION: QEM reduces noise-induced bias in expectation values by post-processing outputs from ensembles of circuit runs at the original or higher noise level.The definition concerns the ensemble-level estimate; individual circuit runs remain noisy.
- I. INTRODUCTION: QEM becomes impractical when accumulated circuit noise completely damages the output, creating a maximum circuit size for a given hardware setup.That size depends on circuit depth and qubit number.
- I. INTRODUCTION: Useful mitigation methods should combine modest qubit overhead, accuracy guarantees, experimental simplicity, and few assumptions about the prepared final state.Formal error bounds should indicate which hardware improvements would improve estimates, while strong final-state assumptions can restrict computationally advantageous applications.
- I. INTRODUCTION: Mitigation trades lower bias against higher variance, so methods can range from quickly converging but less accurate estimators to costlier, more accurate ones.Variance decreases with more circuit runs, whereas systematic bias cannot be reduced by increasing the number of runs.
C. Faults in the circuit
The review models circuit noise as probabilistic faults and uses the circuit fault rate and fault-free probability to quantify mitigation difficulty. QEM can help at moderate noise, but large fault rates incur exponential sampling costs and impose practical circuit-size limits.
- Noise is modeled as discrete probabilistic faults occurring at gates, idling steps, and measurement locations.
- The fault-free probability decays exponentially with the number of fault locations and lower-bounds output-state fidelity under the stated noise assumption.
- The circuit fault rate λ is the average number of faults per run; with M equal-rate locations, λ = Mp, and faults are Poisson-distributed when λ ∼1.
- Fault-free post-selection is unbiased but requires e^λ more circuit repetitions to match the effective shots of a noise-free machine.
- Exponential sampling overhead persists for expectation-value mitigation at large λ and for linear combinations of noisy-circuit outputs.
- QEM avoids active correction but cannot efficiently handle circuits with large λ; its lower implementation cost nevertheless extends near-term device applications.
III. METHODS
The review presents QEM methods that estimate ideal observables from noisy circuit variants, including zero-noise extrapolation and probabilistic error cancellation. These methods can reduce bias, but their assumptions and sampling costs constrain practical scalability.
- III. METHODS: QEM methods construct error-mitigated estimators from noisy primary circuits and circuit variants rather than directly correcting each run.
- A. Zero-noise extrapolation: Zero-noise extrapolation fits noisy expectation values measured at boosted fault rates and evaluates the fitted function at λ = 0.
- A. Zero-noise extrapolation: Richardson extrapolation uses M noisy data points and a degree M −1 polynomial, with linear extrapolation as the M = 2 case.
- A. Zero-noise extrapolation: Three noisy expectation values at λ = 0.5, 1, 1.5 illustrate zero-noise extrapolation, with higher rates produced by boosting device noise.
- A. Zero-noise extrapolation: Richardson extrapolation becomes infeasible when a probed fault rate is too large or data-point gaps are too small; equal-gap overhead scales as (2^M −1)^2.
- A. Zero-noise extrapolation: Richardson extrapolation was effective in a 26-qubit strongly entangled Ising simulation and produced results agreeing with classical simulations for circuits of 127 qubits and 60 two-qubit-gate layers.
- B. Probabilistic error cancellation: Probabilistic error cancellation can fully remove expectation-value bias for generic circuits, while its sampling overhead grows exponentially with circuit fault rate.
- B. Probabilistic error cancellation: Probabilistic error cancellation represents channels as noisy basis operations learned from hardware, then combines them to estimate an ideal channel.
C. Measurement error mitigation
Measurement error mitigation models noisy readout with an assignment matrix and uses inversion, rescaling, or constrained estimation to recover ideal expectations. Scalability, noise correlations, and sampling validity determine the method’s practical limits.
- Measurement errors include SPAM errors, with final readout noise introducing bias into the expectation value of interest.
- The assignment matrix A maps ideal computational-basis outcomes to noisy outcomes, and its entries are estimated from prepared computational states.
- If A is full rank, noisy output distributions can be transformed to obtain ideal expectations once A and the noisy distribution are known.
- Direct assignment-matrix estimation is not efficiently scalable because the matrix dimension grows exponentially with qubit number N.
- Assuming independent measurement errors factorizes A into single-qubit matrices, but realistic experiments can contain correlations that this model misses.
- A continuous-time Markov construction learns 2^N positive coefficients from polynomially many input strings and forms A^-1 = e^-G.
- Pauli twirling diagonalizes the measurement channel in the Pauli basis, reducing mitigation to rescaling selected noisy measurements.
- Matrix inversion may yield negative estimated probabilities under sampling noise, whereas maximum likelihood and Bayesian unfolding enforce valid distributions.
D. Symmetry constraints
Symmetry constraints mitigate errors by rejecting or reweighting outcomes inconsistent with known physical symmetries. Their effectiveness depends on measurable symmetries, while pass rates, measurement overhead, and scalability limit performance.
- D. Symmetry constraints: Post-selection discards runs failing symmetry checks, while post-processing can implement effective projection without simultaneously measuring the symmetry and target observable.
- D. Symmetry constraints: Symmetry verification uses operators commuting with the Hamiltonian, including parity, total spin-Z, and particle-number symmetries.
- D. Symmetry constraints: Symmetry verification can project a circuit’s final state into the appropriate spin or particle-number sector even when the circuit does not conserve that symmetry.
- D. Symmetry constraints: Joint spin-hopping and total-spin measurements require O(N) distinct measurements, whereas hopping terms alone need only two rotation-and-readout choices.
- D. Symmetry constraints: The symmetry projector and symmetrized observable can be decomposed into Pauli operators whose expectation values supply the verified estimate.
- D. Symmetry constraints: Direct verification has sampling overhead Cem ∼ Tr[Πρ]^-1, while post-processing has Cem ∼ Tr[Πρ]^-2.
- D. Symmetry constraints: Symmetry methods cannot mitigate errors commuting with every available symmetry; artificial symmetry schemes can make local operators highly non-local and relatively unscalable.
- D. Symmetry constraints: The MLSC encoding can mitigate or correct all single-qubit errors but does not appear extendable to arbitrarily large code distances.
E. Purity constraints
Purity-based mitigation uses multiple noisy copies or time-separated executions to suppress incoherent errors by emphasizing the dominant eigenstate. These methods can substantially reduce bias, but their sampling and implementation costs limit scalability.
- Virtual distillation (VD) or error suppression by derangement (ESD) uses collective measurements of M noisy copies to estimate observables on the Mth-degree purified state.
- Under p1 > p2, purification converges exponentially in M toward the dominant eigenvector, leaving coherent mismatch as the residual bias.
- Multiple orders of magnitude of error suppression have been confirmed for VD/ESD, even with M = 2 copies, when incoherent noise dominates.
- Transversal operations among copies avoid global measurements but can require challenging long-range interactions; shadow-tomography alternatives scale exponentially in qubit number.
- Echo verification (EV) uses two copies separated in time and post-selects using a noisy projector implemented through the inverse primary circuit.
- EV corrects any noise source with Hermitian Kraus operators to first-order in control-free phase estimation.
- VD/ESD sampling costs grow exponentially with circuit fault rate, while EV cannot be efficiently parallelised even for commuting observables.
F. Subspace expansions
Subspace expansion mitigates errors by optimizing a linear combination of prepared basis states using measured matrix elements. Its success depends critically on choosing a compact, informative basis rather than the exponentially large complete Hilbert-space basis.
- Subspace expansion constructs a target state as a weighted linear combination of easily prepared basis states selected using knowledge of the task.
- The optimal coefficients are obtained by solving a generalized linear eigenvalue problem from Hamiltonian and overlap matrices.
- Improved expectation values are computed from the optimal weights and measured basis-state matrix elements without explicitly preparing the optimized state.
- Choosing the right basis is decisive: a complete Hilbert-space basis returns the ideal state but requires exponentially large classical optimization.
- The initial basis state is usually the best state available before expansion, so the method returns that state in the worst case.
- Symmetry operators can serve as expansion basis operators, allowing subspace expansion to recover the symmetry subspace and incorporate purification-based mitigation.
G. N-representability
N-representability mitigation uses known structural constraints on reduced density matrices (RDMs) to project noisy, invalid estimates back toward representable ones. The approach is especially relevant because low-order RDMs can determine properties of interacting fermionic systems.
- Reduced density matrices are lower-dimensional marginals obtained by integrating out qubits from a joint quantum state.
- Not every marginal is consistent with a valid wave function; the equality and inequality constraints defining valid RDMs comprise the N-representability problem.
- For pairwise-interacting fermions, the 1-RDM and 2-RDM contain the properties of interest, and the 2-RDM determines the energy while remaining exponentially smaller than the full density matrix.
- Useful constraints include Hermiticity, antisymmetry, contraction relations, fixed trace, and positive semidefiniteness.
- The listed constraints are not exhaustive, and pure-state representability can be difficult to apply broadly.
- N-representability mitigation projects noisy RDM estimates that violate known conditions back toward the representable set.
- McWeeny purification restores idempotency but is limited because idempotent fermionic 1-RDMs correspond only to Slater determinants, restricting strongly correlated systems.
H. Learning-based
Learning-based QEM calibrates an error-mitigation function using classically simulable training circuits related to the primary circuit, then applies it to estimate the ideal expectation value. Training can support linear rescaling and shifting, probabilistic error cancellation, and other methods, but the training set and sampling cost can become impractical as system size grows.
- H. Learning-based: Learning-based QEM obtains mitigation parameters by fitting noisy training-circuit results to known ideal expectation values, then applies the learned function to a primary circuit.Training circuits are assumed sufficiently similar to the primary circuit that the learned protocol transfers to it.
- H. Learning-based: When the mitigation function is linear in its parameters, linear least squares can determine the optimal parameters from training data.The same learning framework can use response measurement circuits, including variants with added or replaced gates and measurements.
- H. Learning-based: Linear rescaling and shifting uses noisy expectation values as inputs and can be trained with Clifford variants or problem-specific free-fermion circuits.Under global depolarising noise, the linear form follows from the noisy state being a mixture of the ideal and maximally mixed states.
- H. Learning-based: Training circuits should resemble the primary circuit and preserve similar faults, while remaining classically simulable so their ideal expectation values are known.For probabilistic error cancellation, Clifford variants can replace gates while retaining comparable errors.
- H. Learning-based: When the Clifford training loss reaches zero, the learned function exactly recovers ideal expectation values across the targeted unitary circuit family.This ideal result is stated for all circuits in the family, including the primary circuit.
- H. Learning-based: The Clifford training set grows exponentially with qubit number, so practical implementations truncate it, sample circuits, or retain only selected non-Clifford gates.Training cost is problem-specific, and larger training sets can shift sampling overhead to device calibration rather than task execution.
A. Comparison among QEM methods
QEM methods can be compared by separating noise calibration from response measurement, but their costs and performance depend strongly on the method, circuit, assumptions, and hardware constraints. The review therefore uses canonical implementations to summarize distinguishing features rather than claiming universal rankings.
- A. Comparison among QEM methods: QEM consists of noise calibration, which measures noise strength, followed by response measurement, which measures observable changes under calibrated noise.Together, these stages support construction of an estimator protected against the calibrated noise components.
- A. Comparison among QEM methods: Gate-error methods can calibrate through standard benchmarking before QEM application, but correlated noise characterization is exponentially expensive and device drift may require recalibration.Calibration is effectively free during application only when completed during device calibration and parameters remain stable.
- A. Comparison among QEM methods: State-error methods calibrate by measuring violations of known output-state constraints, such as symmetry or purity conditions.Observable-error methods instead target error components damaging to a selected observable, making calibration accuracy and cost problem-dependent.
- A. Comparison among QEM methods: Learning-based methods use training circuits related to the primary circuit to target more specific faults and potentially reduce calibration cost, but training cost remains problem-specific.Their training cost is difficult to quantify analytically, similar to observable-error mitigation calibration.
- A. Comparison among QEM methods: QEM categories are heuristic rather than definitive because a method such as N-representability can be interpreted as observable-error or state-error mitigation.The interpretation depends on whether the reduced density matrix is viewed as an observable or a state in a subspace.
- A. Comparison among QEM methods: General method comparisons require a specified primary circuit and experimental constraints because estimators trade bias against variance and methods use different assumptions and hardware resources.The review therefore summarizes canonical implementations in Table IV rather than asserting universal performance orderings.
1. State discrimination
State-discrimination analyses frame QEM as extracting or distinguishing ideal components from noisy states and derive limits on estimator range, bias, and sampling cost. These results show that noise and circuit size can force exponential sampling overhead, while extraction rate provides a cost-effectiveness measure.
- 1. State discrimination: QEM cannot increase distinguishability between noisy output states, which bounds the range of error-mitigated estimators and the samples needed for a target performance.These bounds describe fundamental limits for a given setup rather than practical metrics for ranking methods.
- 1. State discrimination: The maximum QEM bias is bounded through state discrimination, and trace distance supplies a lower bound on the estimator range used to determine sufficient sampling.The range connects state distinguishability to the number of samples required for estimation.
- 1. State discrimination: Explicit bounds relate the necessary sample count to target accuracy and success probability through relative entropy between noisy states.Relative entropy is connected to trace distance by quantum Pinsker’s inequality.
- 1. State discrimination: Exponential decay of relative entropy under local depolarising noise makes the sampling cost grow exponentially with circuit depth.Related constructions show worst-case sampling lower bounds can scale exponentially with both circuit depth and qubit number.
- 1. State discrimination: For layered circuits under broad Markovian noise, quantum Fisher information decays exponentially with depth, implying exponentially increasing QEM sampling overhead.Under global depolarising noise, rescaling the measurement result saturates the corresponding sampling bound.
- 1. State discrimination: The extraction rate is the fidelity boost divided by the square root of sampling overhead and measures QEM cost-effectiveness.Direct symmetry verification can achieve extraction rate 1 by extracting all error-mitigated components, whereas nonlinear response combinations fall outside this analysis.
C. Combinations of QEM methods
QEM methods can be combined in parallel, concatenated sequentially, or interpolated through unified parameterised frameworks. Their interactions can improve coverage of noise sources but may multiply sampling costs or alter the assumptions required by later stages.
- C. Combinations of QEM methods: Parallel combinations assign different QEM methods to different noise sources, such as circuit noise and measurement noise.Zero-noise extrapolation or probabilistic error cancellation can target computation while measurement-error mitigation targets readout.
- C. Combinations of QEM methods: When QEM methods interfere, application order matters, and concatenating two stages multiplies their sampling overheads.Symmetry or purification methods can be applied after treating a base method as extracting an error-mitigated state.
- C. Combinations of QEM methods: Applying zero-noise extrapolation after another method may invalidate the noise-scaling factor, while probabilistic error cancellation becomes difficult when the effective error model changes.These concerns arise because the base method can alter the circuit’s error model.
- C. Combinations of QEM methods: Learning-based methods combine readily with other QEM methods because they can replace their noise-calibration procedures or optimise their hyperparameters.In this role, learning-based QEM functions as calibration rather than an isolated mitigation stage.
- C. Combinations of QEM methods: Unified parameterised frameworks interpolate between existing QEM methods and expose additional implementations through intermediate hyperparameter values.This interpolation differs from concatenating separate mitigation stages.
- C. Combinations of QEM methods: QEC and QEM are often complementary: QEM can be applied on top of QEC and can mitigate compilation errors that QEC cannot address.Symmetry verification generally needs fewer qubits and lower gate fidelity than QEC to reach an unmitigated-circuit break-even point, though exact differences depend on the application.
V. APPLICATIONS
The review surveys QEM applications across hardware and logical settings, emphasizing noise-type-dependent strategies and resource trade-offs. It reports experimental use across diverse applications and discusses extending mitigation to fault-tolerant computation.
- QEM software packages implement zero-noise extrapolation, probabilistic error cancellation, and measurement error mitigation.
- QEM has been effective in numerical studies and physical experiments spanning linear equation solvers, metrology, Monte Carlo simulations, and Fermi-Hubbard models.
- A. Coherent errors: Pauli noise is generally easier to analyse and mitigate than coherent noise, while coherent errors often accumulate faster.
- A. Coherent errors: Pauli twirling transforms arbitrary quantum channels into Pauli channels by inserting random Pauli gates during circuit compilation.
- B. Logical errors in fault-tolerant quantum computation: For logical circuits, probabilistic error cancellation can reduce physical-qubit overhead by 80% for some classically intractable problems at fault rate 10^-3 and sampling overhead 100.
- B. Logical errors in fault-tolerant quantum computation: With perfect Clifford gates, probabilistic error cancellation can remove errors from circuits with 2000 T gates at physical error rate 10^-3 and sampling overhead 1000.
C. Algorithmic (compilation) errors
The review treats compilation and other algorithmic errors as distinct from physical noise and surveys QEM methods that can mitigate them. It also identifies unresolved questions about classification, metrics, optimization, fault-tolerant connections, and scope.
- C. Algorithmic (compilation) errors: Compilation mismatches and imperfect variational ansatz parameters produce algorithmic errors even when gates execute without physical noise.
- C. Algorithmic (compilation) errors: Unlike quantum error correction, some QEM techniques can remove algorithmic errors because they arise from imperfect computation rather than noisy gate execution.
- C. Algorithmic (compilation) errors: Zero-noise extrapolation can mitigate algorithmic errors by combining circuits with different structures or parameters when their relative error scaling is known.
- C. Algorithmic (compilation) errors: Randomised Trotterisation provides another setting where constraint-based QEM can address algorithmic errors that violate known properties of the ideal state.
- Open questions concern better QEM classifications, practical performance metrics, optimal method combinations, and systematic connections with quantum error correction.
- The review also asks whether QEM can extend beyond expectation-value estimation to repeat-until-success and single-shot algorithms.
Appendix A: Practical Techniques in Implementations
The appendix compares practical estimators for QEM, focusing on how linear combinations can be sampled efficiently and how noise affects sampling overhead. Monte Carlo sampling avoids the scalability problem of estimating many components separately, but overhead can still grow exponentially with noise.
- QEM estimators often express the mitigated expectation value as a linear combination of K response-circuit observables.
- 1. Monte carlo sampling: Estimating each component separately is not scalable when K is large, motivating a probabilistic-mixture Monte Carlo estimator.
- 1. Monte carlo sampling: The Monte Carlo estimator rescales samples from a mixture of signed component observables by A, where mixture probabilities are proportional to coefficient magnitudes.
- 1. Monte carlo sampling: Under the stated assumptions, Monte Carlo sampling is always more sample efficient than separately estimating and combining the component observables.
- a. Exponential sampling overhead: For Pauli observables under Pauli circuit noise, the mixture variance becomes comparable to component variance at large noise levels.
- a. Exponential sampling overhead: Sampling cost increases exponentially with the circuit fault rate λ when mitigating Pauli observables under Pauli circuit noise.
2. Pauli twirling
Pauli twirling standardizes noise by removing off-diagonal channel components, making Pauli-noise analysis and measurement procedures central to the reviewed techniques. The section also describes circuits for measuring Pauli and product observables.
- 2. Pauli twirling: Twirling a noise process over a symmetry group means conjugating it with random group elements.
- 2. Pauli twirling: Pauli twirling removes off-diagonal elements in the Pauli basis and transforms an arbitrary error channel into Pauli noise.
- 2. Pauli twirling: For noisy Clifford gates, random Pauli gates and corresponding conjugated Paulis can be inserted in each circuit run to twirl the gate noise.
- 3. Measurement techniques: The Hadamard test measures Re(Tr[Oρ]) through X⊗I and Im(Tr[Oρ]) through Y⊗I for a general unitary O.
- 3. Measurement techniques: Arbitrary Pauli operators can be measured by single-qubit Pauli measurements followed by multiplying component outcomes, requiring only one layer of basis-changing Clifford gates.
- 3. Measurement techniques: Two commuting Pauli operators can share a circuit run with this post-processing approach only when they qubit-wise commute.
- 3. Measurement techniques: For a Pauli O and Hermitian S, the displayed circuits measure the commuting component S+; the anticommuting component requires an additional Pauli rotation.