Source-linked AI summary
Virtual Distillation for Quantum Error Mitigation
William J. Huggins, Sam McArdle, Thomas E. O'Brien, Joonho Lee, Nicholas C. Rubin, Sergio Boixo, K. Birgitta Whaley, Ryan Babbush, Jarrod R. McClean
TL;DR
Noisy near-term quantum computers need error-mitigation strategies before fault-tolerant correction is practical. The paper introduces virtual distillation, which measures M noisy copies to access ρ^M/Tr(ρ^M) without preparing it, and reports strong error suppression while identifying hardware and noise constraints.
Problem
Near-term quantum computers have high error rates, while the overhead of fault-tolerant quantum error correction is currently impractical.
Method
Virtual distillation uses collective measurements of M copies of a noisy state to estimate observables of ρ^M/Tr(ρ^M) without explicitly preparing the purified state.
Results
Numerical experiments demonstrate trace-distance error reductions of up to three orders of magnitude, enhanced with system size or information-scrambling speed.
Takeaways & Limitations
The strategy can use surplus qubits to mitigate incoherent device and algorithmic errors, with the effective state approaching ρ's dominant eigenvector exponentially in M.
Takeaways & Limitations
Performance requires collective measurements coupling corresponding qubits across copies, which can require substantial additional gates on general 2D hardware connectivity.
Abstract
from arXiv · showhide
Contemporary quantum computers have relatively high levels of noise, making it difficult to use them to perform useful calculations, even with a large number of qubits. Quantum error correction is expected to eventually enable fault-tolerant quantum computation at large scales, but until then it will be necessary to use alternative strategies to mitigate the impact of errors. We propose a near-term friendly strategy to mitigate errors by entangling and measuring $M$ copies of a noisy state $ρ$. This enables us to estimate expectation values with respect to a state with dramatically reduced error, $ρ^M/ \mathrm{Tr}(ρ^M)$, without explicitly preparing it, hence the name "virtual distillation". As $M$ increases, this state approaches the closest pure state to $ρ$, exponentially quickly. We analyze the effectiveness of virtual distillation and find that it is governed in many regimes by the behavior of this pure state (corresponding to the dominant eigenvector of $ρ$). We numerically demonstrate that virtual distillation is capable of suppressing errors by multiple orders of magnitude and explain how this effect is enhanced as the system size grows. Finally, we show that this technique can improve the convergence of randomized quantum algorithms, even in the absence of device noise.
I. INTRODUCTION
Virtual distillation uses collective measurements on multiple noisy copies to estimate observables of an effectively purified state without explicitly preparing that state. The approach can exploit extra qubits and suppress nondominant eigenvector contributions exponentially with the number of copies.
- Virtual distillation estimates expectation values for ρ^M/Tr(ρ^M) using collective measurements of M copies, rather than explicitly preparing a purified state.
- The effective state's nondominant eigenvector weights are suppressed exponentially in M, compared with linear suppression for explicit approximate purification.
- The strategy aims to use additional qubits to improve noisy computations without the large overhead of traditional quantum error correction.
- The approach can be limited to a constant-factor error-rate improvement in the worst case, and purely coherent errors receive no protection.
- The method assumes separate copies experience noise with the same form and strength; unentangled unequal-noise copies instead produce an effective product of their density matrices.
- Its basic implementation uses repeated pairs of copies, applies two-qubit gates between corresponding qubits, and measures computational-basis outcomes.
A. Measurement by Diagonalization
For single-qubit observables with two copies, virtual distillation can be implemented by diagonalizing the relevant swap-observable products with pairwise two-qubit gates. This yields simultaneous computational-basis estimates for all corresponding single-site observables.
- For M = 2 and a single-qubit Pauli Z observable, pairwise two-qubit operations provide a straightforward virtual-distillation implementation.
- The implementation relies on diagonalizing products of the swap operator and pairwise observables with a factorized two-qubit unitary.
- A single layer of N two-qubit gates followed by computational-basis measurement estimates the numerator and denominator needed for corrected expectation values.
- The same measurement collects error-mitigated expectation values for all N single-qubit Z operators simultaneously.
- Appropriate single-qubit rotations extend the procedure from Z to arbitrary single-qubit observables.
B. Sample Efficiency
The paper analyzes the sampling cost of virtual distillation and proposes collective measurements that can outperform naive independent repetition. Under specified conditions, the improved estimator's variance can decrease quadratically with the number of measurement groups.
- The required circuit repetitions depend on estimator variance, which increases as the purity of ρ decreases.
- The collective-measurement generalization is motivated by settings with sufficiently high noise and access to many more copies than the mitigation order requires.
- The proposed collective measurement uses all 2K copies when M = 2 and can outperform running K independent protocol pairs in parallel.
- When Tr(ρ^3) is small, the collective estimator's variance shrinks quadratically with K in an identified regime, versus linear suppression for naive parallelization.
- The paper does not propose a specific NISQ-friendly implementation of the collective operator or establish comprehensive limits on virtual-distillation sample complexity.
III. PERFORMANCE UNDER DIFFERENT NOISE MODELS
Virtual distillation suppresses errors rapidly when errors move the state into orthogonal sectors, while estimator variance and dominant-eigenvector behavior determine its practical performance.
- Orthogonal Errors: Orthogonal-error models leave the dominant eigenvector equal to the ideal state, allowing virtual distillation to remove error weight through repeated copies.The analysis models stochastic errors as transitions into new orthogonal states.
- Orthogonal Errors: Quadratic suppression of errors is expected in the most favorable M = 2 case.For M copies, the relevant density-matrix coefficients are raised to the Mth power.
- Sample Complexity: Estimator variance depends on the sample complexity of estimating the corrected numerator and denominator from independent experiments.For M = 2, the analysis uses R independent experiments for each quantity, totaling 2R experiments.
- Sample Complexity: At fixed additive error, the required number of measurements can be used to estimate the corrected expectation value.The supplied passage states this estimation guarantee without specifying the measurement count.
B. Non-Orthogonal Error Floor
Non-orthogonal errors shift the density matrix’s dominant eigenvector away from the ideal state, creating an error floor that generally prevents quadratic low-noise suppression.
- Non-Orthogonal Error Floor: Non-orthogonal error population can drift the dominant eigenvector away from the target state, limiting virtual distillation’s maximum improvement.This drift is the practical counterpart to the idealized assumption that the dominant eigenvector is exactly the noiseless state.
- Perturbative Analysis: In the low-error regime, matrix perturbation theory expands the dominant eigenvector as a convergent power series in the perturbation strength.The analysis assumes λ0 ≫ λ1 and ∆ ≪ |λ0 − λ1|.
- Non-Orthogonal Error Floor: Without further constraints on the state, noise model, or observables, low-noise behavior is generally limited to a constant-factor improvement rather than quadratic suppression.The improvement magnitude depends on the typical size of γ.
- Exceptions: Quadratic suppression can reappear at intermediate error rates or for cases where γ vanishes or particular observables evade the trace-distance floor.The simulations report behavior consistent with quadratic suppression at intermediate error rates.
- Exceptions: Symmetries and strongly scrambling dynamics can produce conditions under which γ is near zero.The passage gives symmetry-group errors and certain chaotic circuits as examples.
IV. NUMERICAL EXPERIMENTS
The numerical study evaluates virtual distillation on random circuits and one-dimensional spin-chain quenches, using trace distance to compare the effective and ideal states.
- Systems: The simulations cover three model systems: two classes of random circuits and a one-dimensional spin chain after a quantum quench.The random-circuit cases provide analytically tractable limits, while the quench studies probe quantum-simulation dynamics.
- Systems: The spin-chain study focuses on time evolution rather than ground states because ground states have additional structure supporting specialized error-mitigation techniques.This choice examines virtual distillation without that additional structure.
- Noise Model: Performance is characterized by expected gate-error count, enabling comparisons that trade off per-gate error rate against circuit depth.The noise model applies single-qubit depolarizing channels to both qubits after each two-qubit gate.
- Evaluation: Trace distance between the ideal noiseless state and the virtual-distillation effective state is the main error metric.Trace distance supplies a bound on expectation-value error for arbitrary observables.
A. Scrambling Circuits
Virtual distillation suppresses errors in random circuits by using multiple noisy copies, but entangling circuits retain a floor set by drift in the dominant eigenvector. This floor generally decreases with system size, while non-entangling circuits exhibit exponential suppression with copy number.
- Non-entangling circuits: Non-entangling circuits show nearly linear error behavior for M = 1, 2, and 3 at low expected error counts, with slopes 1, 2, and 3 respectively.Their eigenvalue floor vanishes, producing exponential suppression as the number of copies increases.
- Entangling circuits: Entangling circuits have a nonzero dominant-eigenvector error that sets the minimum error achievable by virtual distillation, independent of M.The floor grows slowly with circuit depth and is suppressed as system size increases.
- Entangling circuits: For fixed circuits at low error rates, dominant-eigenvector errors scale linearly with the error rate and are orders of magnitude smaller than unmitigated-state errors.The resulting floor is also suppressed as system size increases.
- Heisenberg quench: In Heisenberg-model simulations, virtual distillation suppresses actual single-site magnetization errors similarly to its suppression of trace distance.Trace-distance bounds are roughly an order of magnitude too pessimistic for this observable, while noisy and noiseless two-copy distillation nearly coincide.
- Heisenberg quench: For Heisenberg evolution, increasing system size from 6 to 10 qubits decreases trace-distance error for error-mitigated states with M > 1.The improvement factor appears substantial but is smaller and less system-size-sensitive than for one-dimensional scrambling circuits.
- Practical cost: Virtual distillation requires only a modest measurement-cost increase when expected errors are small, but the cost rises dramatically once expected errors exceed one.This identifies a practical operating regime for the technique.
V. MITIGATING ALGORITHMIC ERRORS
Virtual distillation can mitigate incoherent algorithmic errors in randomized evolution, including qDRIFT, by reducing the coherent resources needed to reach a target accuracy. In the studied Heisenberg model, the method produced an 8x space-time advantage after accounting for two-copy overhead.
- V. MITIGATING ALGORITHMIC ERRORS: Virtual distillation suppresses incoherent algorithmic deviation in randomized evolution, including qDRIFT, even without device noise.Randomized methods output mixed states and converge toward exact evolution as their approximation parameter increases.
- V. MITIGATING ALGORITHMIC ERRORS: 16x fewer coherent qDRIFT steps were required to reach trace distance 0.01 in the studied Heisenberg model.The simulations used Heisenberg Hamiltonians with up to 6 qubits per copy and evolution length t = N.
- V. MITIGATING ALGORITHMIC ERRORS: 8x lower coherent space-time volume remained after accounting for the overhead of using two copies.This comparison is for reaching the same target error rate in the Heisenberg-model study.
- V. MITIGATING ALGORITHMIC ERRORS: The broader study found error reductions of up to three orders of magnitude and stronger suppression as system size or information-scrambling speed increased.The qDRIFT application was reported as a substantial constant-factor improvement.
- V. MITIGATING ALGORITHMIC ERRORS: Virtual distillation is expected to be most effective and affordable when the expected number of circuit errors is O(1).The authors associate its growing reach during the NISQ era with improving hardware error rates.
- V. MITIGATING ALGORITHMIC ERRORS: Collective measurements and multi-qubit observables introduce connectivity, gate-complexity, and measurement-repetition overheads.These practical considerations affect the technique’s performance on near-term hardware.
Appendix A: Measuring Multi-Qubit Observables by Diagonalization
Diagonalization-based virtual distillation extends from two-copy measurements to higher powers and multi-qubit observables, but general multi-qubit measurements remain operationally constrained. For three copies, a numerically optimized ansatz approximately diagonalizes the required cyclic-shift structure.
- Appendix A: Measuring Multi-Qubit Observables by Diagonalization: Multi-qubit observables can be handled using a nonsymmetrized form that avoids ancilla-assisted measurement and circuit depth.The alternative is motivated by the failure of the symmetrized observable to factorize across corresponding qubit pairs.
- Appendix A: Measuring Multi-Qubit Observables by Diagonalization: Noncommuting measurement structure prevents simultaneous estimation of some numerator and denominator terms, while commuting multi-qubit observables may still require separate measurements.This can increase the total number of circuit repetitions.
- Appendix B: Measurement by Diagonalization with Three or More Copies: A diagonalization circuit can factorize across qubit tuples when the relevant shift and symmetrized operators factorize in the same way.For M copies, the cyclic shift decomposes into single-qubit cyclic shifts across the N tuples.
- Appendix B: Measurement by Diagonalization with Three or More Copies: Diagonalization-based protocols estimate Tr(Z_kρ^M) and Tr(ρ^M), but reconstructing a multi-qubit Pauli expectation requires measuring one particular operator at a time.The same concern applies to the generalized higher-copy proposal.
- Appendix B: Measurement by Diagonalization with Three or More Copies: For M = 3, the optimized four-gate ansatz achieved approximately 5E−5 Frobenius-norm error relative to the exact diagonalization.The ansatz is shown in Figure 11 and was numerically optimized to approximately diagonalize S(3).
Appendix C: Ancilla-Assisted Measurement
Ancilla-assisted virtual distillation uses controlled swap or cyclic-shift tests to estimate corrected expectation values from multiple copies. Its collective-measurement construction can support broader observable measurements, while a related joint measurement can reduce variance under certain conditions.
- Appendix C: Ancilla-Assisted Measurement: The ancilla-assisted protocol performs a nondestructive swap measurement on two copies and an ancilla, then measures the observable on the system registers.A Hadamard test estimates the swap expectation, and the symmetrized observable is measured afterward.
- Appendix C: Ancilla-Assisted Measurement: Unlike the diagonalization approach, the ancilla-assisted variant supports arbitrary measured operators and simultaneous measurement on overlapping qubit subsets.The method is therefore compatible with techniques for efficiently measuring large collections of commuting operators.
- Appendix C: Ancilla-Assisted Measurement: For three or more copies, a controlled cyclic-shift test estimates Re(Tr(S(N)ρ⊗N)), followed by measurement of the commuting symmetrized observable.The resulting expectation is Tr(S(N)O(N)ρ⊗N).
- Appendix E: Variance of the Proposed Collective Measurement: The variance analysis estimates numerator and denominator moments and their covariance from repeated measurements, using a Taylor approximation for the ratio.The approximation is expected to be good when the number of repetitions R is sufficiently large.
- Appendix E: Variance of the Proposed Collective Measurement: A joint measurement on 2K copies can estimate Tr(Oρ^2) with lower variance than K parallel basic virtual-distillation procedures under certain conditions.Appendix E proves the claimed variance comparison for the collective operator.
Appendix F: Details Regarding the Numerical Experiments with Scrambling Circuits
This appendix details scrambling-circuit simulations and analyzes how depolarizing noise shapes the final density matrix and its eigenvectors. For submaximal noise, the dominant eigenvector remains the ideal state, while subsequent eigenvectors represent increasingly many qubit errors.
- Circuit construction: The scrambling circuits alternate two-qubit Sycamore-gate layers with randomly chosen single-qubit gates.Two-qubit layers alternate even-odd and odd-even pairings, with random single-qubit layers inserted between them.
- Circuit construction: A second circuit class removes the two-qubit gates while retaining single-qubit depolarizing noise at the same locations.This isolates the effects of single-qubit gates and noise from entangling dynamics.
- Noise analysis: The final density matrix is expressed analytically using noiseless single-qubit states, depolarizing probability p, and circuit depth D.Repeated channel applications are reduced to an equivalent single application with an effective error rate.
- Eigenvector structure: For any p below its maximum value, the density matrix’s dominant eigenvector is exactly the ideal product state.When p is small, the next-largest eigenvectors correspond to states with single-qubit errors, followed by states with multiple errors.
Appendix G: Interplay with the surface code
This appendix compares virtual distillation with surface-code error correction when additional qubits are available. Although surface-code errors improve exponentially with code distance, overhead may make virtual distillation’s constant-factor improvement advantageous at moderate distances.
- Surface-code overhead: Surface-code resource requirements grow with code distance because physical qubits scale as n = 2d^2 and operation cycles scale with d.Arbitrary rotations can require approximately 200d surface-code cycles under the stated synthesis heuristic.
- Resource comparison: The comparison uses twice the qubits either for virtual distillation or for increasing the surface-code distance.The appendix frames this as a fixed-resource comparison between virtual distillation and stronger logical protection.
- Error suppression: Virtual distillation provides a large constant-factor improvement over the bare circuit, whereas surface-code error correction offers exponential returns asymptotically.The relevant practical question is whether surface-code overhead outweighs virtual distillation’s constant-factor benefit at finite distances.
- Finite-distance regime: Up to distance 10−15, the appendix reports that the empirical improvement can make virtual distillation advantageous under some conditions.The analysis is approximate and treats code distances continuously rather than rounding them to integers.
Appendix H: Virtual Distillation Applied to Distinct States
This appendix studies virtual distillation when the two input copies experience different noise and therefore contain distinct states. It identifies a physicality limitation and tests the approach using bit-flip and phase-flip noise.
- Distinct-state extension: The analysis extends virtual distillation from identical copies to two distinct noisy states, ρA and ρB.The effective expectation values are derived for collective measurements applied to these different states.
- Assumption: The derivation assumes that the two copies are unentangled before virtual distillation.A diagrammatic proof is used to establish the resulting effective measurement expression.
- Limitation: For distinct states, the effective matrix ρeff is not guaranteed to be positive semidefinite or represent a valid quantum state.This requires care in variational algorithms to avoid obtaining a non-physical, non-variational result.
- Numerical test: The numerical test combines bit-flip noise for ρA with phase-flip noise for ρB in the Heisenberg evolution.The resulting effective-state error is examined in Figure 13.
Appendix I: Performance Under Amplitude Damping and Dephasing Noise
This appendix evaluates virtual distillation under amplitude-damping and dephasing noise, using random circuits and Heisenberg evolution. The qualitative benefits persist, although the magnitude and system-size dependence differ across circuit families.
- Noise model: The realistic noise model combines amplitude damping, parameterized by γ1, with dephasing, parameterized by γ2.Amplitude damping describes decay from |1⟩ to |0⟩, while dephasing models environment-induced computational-basis measurement.
- Evaluation: The simulations quantify error by trace distance to the noiseless state as a function of the expected number of errors.The parameters are chosen so the associated T1 and T2 times are comparable.
- Simulation design: Figures 14 and 15 compare unmitigated states, virtual-distillation states for M = 2 and 3, and dominant eigenvectors for random circuits and Heisenberg evolution.The figures use the same respective circuit families studied in the main text.
- Results: For sufficiently small expected error counts, virtual distillation still suppresses error by orders of magnitude, with M = 2 often sufficient.The performance remains limited by drift of the dominant eigenvector.
- Results: In ten-qubit random circuits, the benefit is roughly two orders of magnitude instead of three, while system-size dependence mostly vanishes.For Heisenberg evolution, system-size dependence persists under this noise model.