Source-linked AI summary
Demonstration of qubit operations below a rigorous fault tolerance threshold with gate set tomography
Robin Blume-Kohout, John King Gamble, Erik Nielsen, Kenneth Rudinger, Jonathan Mizrahi, Kevin Fortier, Peter Maunz
TL;DR
Fault-tolerant quantum error correction needs physical operations below diamond-norm thresholds, but randomized benchmarking does not reliably characterize that worst-case error for general noise. This paper applies gate set tomography to a trapped-Yb+ ion qubit and demonstrates all three logic gates below the rigorous threshold with greater than 95% confidence. The result establishes a method for directly verifying threshold compliance while revealing and reducing non-Markovian effects.
Problem
Randomized benchmarking reports an error rate that is not sensitive to all errors and cannot be directly compared with diamond-norm thresholds required for general-noise FTQEC.
Method
Gate set tomography fully characterizes a trapped-Yb+ qubit, supplies confidence bounds, and guides iterative debugging and improvement of its operations.
Results
(1.58 ± 0.15) × 10^-4, (1.39 ± 0.22) × 10^-4, and (1.62 ± 0.27) × 10^-4 diamond norm errors were obtained for GI, GX, and GY, respectively, all surpassing the threshold with 95% confidence.
Takeaways & Limitations
GST provides the first demonstrated single-qubit gates below a rigorous fault-tolerance threshold against general noise and reliable feedback for improving those gates.
Abstract
from arXiv · showhide
Quantum information processors promise fast algorithms for problems inaccessible to classical computers. But since qubits are noisy and error-prone, they will depend on fault-tolerant quantum error correction (FTQEC) to compute reliably. Quantum error correction can protect against general noise if -- and only if -- the error in each physical qubit operation is smaller than a certain threshold. The threshold for general errors is quantified by their diamond norm. Until now, qubits have been assessed primarily by randomized benchmarking, which reports a different "error rate" that is not sensitive to all errors, and cannot be compared directly to diamond norm thresholds. Here we use gate set tomography (GST) to completely characterize operations on a trapped-Yb$^+$-ion qubit and demonstrate with very high ($>95\%$) confidence that they satisfy a rigorous threshold for FTQEC (diamond norm $\leq6.7\times10^{-4}$).
I. INTRODUCTION
Fault-tolerant quantum error correction requires physical operations below noise-model-dependent thresholds, but randomized benchmarking cannot directly establish diamond-norm compliance for general errors. The paper uses gate set tomography to characterize, debug, and improve a trapped-Yb+ qubit, demonstrating all three operations below a rigorous FTQEC threshold.
- Diamond norm error quantifies thresholds against realistic general errors, including small unitary errors.
- Randomized benchmarking primarily estimates average process infidelity and is relatively insensitive to unitary errors, limiting its ability to bound diamond-norm error.
- GST provides a full tomographic description of every gate with statistical confidence bounds, enabling iterative improvement and tight diamond-norm bounds.
- GST is self-calibrating and avoids the calibration-error propagation affecting process tomography while addressing randomized benchmarking’s worst-case-error limitation.
- The method assumes a two-dimensional qubit and stationary, Markovian gate operations, while also detecting and quantifying violations of those assumptions.
B. Experiment
The experiment used GST to characterize and progressively improve three microwave-driven gates on a single trapped 171Yb+ ion. Across five runs, stabilization, drift control, dynamical decoupling, and calibration changes reduced coherent and non-Markovian errors to sub-threshold levels.
- The qubit was a single 171Yb+ ion encoded in the hyperfine clock states |0⟩ and |1⟩.
- March 2015 experiments surpassed the 6.7 × 10^-4 diamond-norm threshold with 95% confidence.
- The implemented gate set comprised GI, the idle gate, plus GX and GY, π/2 rotations about X and Y driven by microwave pulses.
- GST estimates used gate sequences extending to length 8192, with process matrices and error generators compared against target operations.
- Five experimental runs from 17 April 2014 to 30 March 2015 tracked steady improvement in process infidelities.
- Drift control, trap stabilization, dynamical decoupling, and improved BB1 calibration progressively reduced coherent and non-Markovian errors.
C. Demonstrating suitability for fault tolerance
Fault-tolerance suitability is assessed using diamond-norm bounds and tests of Markovianity. GST directly bounds gate errors, while goodness-of-fit analysis evaluates detectable non-Markovian behavior.
- Diamond-norm criterion: Diamond-norm thresholds provide the relevant criterion for fault tolerance against general quantum errors.Threshold theorems are stated using the diamond-norm distance between real and ideal gates.
- GST characterization: GST enables direct diamond-norm computation between estimated and target gates using a semidefinite program.This supports quantitative confidence bounds on the gate errors rather than relying only on randomized benchmarking.
- Scope: The demonstration covers single-qubit operations only and does not itself demonstrate FTQEC, which also requires two-qubit gates, repeatable measurements, and more qubits.The GST methods generalize to two-qubit gates, repeatable measurements, and crosstalk characterization.
- Non-Markovianity: GST’s Markovian model assumes stationary operations, but experimental non-Markovian effects can arise from drift, correlated errors, and leakage.In the presence of non-Markovian noise, GST guarantees are formally void, although many typical violations produce quantifiable fit failure.
- Non-Markovianity: 2∆log L evaluates goodness-of-fit by comparing observed data with the best Markovian gate-set model across gate-sequence collections.Under the stated assumptions, its asymptotic distribution is χ2_k with k = Ns − Np and expected value 2k.
- Non-Markovianity: The March 30 dataset resembled simulated Markovian data, with Markovianity violated at 4σ but judged practically insignificant given the experiment’s sensitivity.Reducing sensitivity by a factor of four would make the observed non-Markovianity undetectable while increasing diamond-norm error bars to ±8 × 10^-5.
E. Comparison to randomized benchmarking
The paper compares GST with randomized benchmarking (RB), showing that similar RB decay rates can mask different underlying gate errors. A composite non-Markovian model consistent with the data helps explain the discrepancy, while GST and RB have different sensitivities to coherent and time-dependent noise.
- RB comparison: (5.31 ± 0.16) · 10^-5 per elementary gate is the experimentally observed RB error rate, while GST predicts (4.53±0.25)·10^-5.The experimental rate is obtained by fitting probabilities versus the number of elementary gates; GST prediction uncertainties come from parametric bootstrap.
- Noise interpretation: GST typically reports higher Markovian noise than RB under low-frequency drift, because GST amplifies coherent errors that RB sequences are relatively insensitive to.In this experiment, however, GST underestimates the RB error rate; the observed opposite effect is consistent with anti-correlated noise from dynamically corrected gates.
- Noise interpretation: Dynamically corrected gates can produce plausible errors that flip sign every clock cycle, because pulse timing or amplitude errors persist in the toggling frame.The dynamically corrected idle gate uses periodic Xπ pulses to echo away small Z rotations, while residual pulse errors can alternate between applications.
- Model comparison: GST cannot distinguish Gcomp from G0: every gate is within 4.4 · 10^-5 in diamond norm of the corresponding gate in G0.The simulation places nearly all free gate-matrix elements within G0’s 95% confidence intervals, with the remaining deviations no more than 0.05σ outside them.
- RB comparison: RB error rates of (5.38 ± 0.17) · 10^-5 for simulated Gcomp and (5.31 ± 0.16) · 10^-5 experimentally are statistically indistinguishable.Gcomp is a composite model with conditional gate sets and alternating coherent over- and under-rotations, yet its simulated RB data matches the experiment almost perfectly.
- Caveat: Neither GST nor RB is designed to function reliably under any non-Markovian noise, so Gcomp is a plausible consistent model rather than a demonstrated physical description.The paper notes that multiple distinct non-Markovian models may be equally consistent with the data.
F. The relative power of RB and GST
RB and GST use repeated long gate sequences but differ fundamentally: RB randomizes to average errors, whereas GST uses structured periodic sequences to amplify them. This makes GST more efficient for detecting coherent errors and bounding diamond norm error, while RB can miss them.
- RB randomizes gate sequences to twirl noise, whereas GST uses structured periodic sequences to amplify errors.
- A coherent rotation θ appears in RB as an incoherent error probability Lθ^2, but periodic circuits can accumulate the rotation to produce L^2θ^2.
- The diamond norm captures worst-case failure growth, scaling as O(θ) for small coherent errors, while process infidelity scales as O(θ^2).
- GST detects coherent errors using sequences of length O(1/θ), whereas randomized sequences require length O(1/θ^2) or many more repetitions.
- Unitarity benchmarking can separate coherent and incoherent errors and inform diamond norm rates, but it is extremely inefficient compared with GST.
- For r = 10^-4, bounding diamond norm error through RB and unitarity requires measuring both quantities to 10^-8 precision.This would require approximately N = 10^8 repetitions, at least 10^6 times more than standard RB or GST, making the approach impractical.
G. Validating 10−5 accuracy with simulations
Simulated GST experiments with known unitary errors validate the claimed diamond-norm precision. Estimation error follows 1/L scaling up to sequence lengths set by the stochastic error rate.
- The 1/L scaling holds up to L ≈ 1/ϵ, where ϵ is the stochastic error rate.
III. DISCUSSION
The discussion places the single-qubit GST result in context: it establishes rigorous threshold performance while remaining only one step toward full fault-tolerant quantum error correction. GST also offers a route to characterize broader multiqubit requirements.
- GST enabled trapped-Yb+-ion gates to become the first trapped-ion gates demonstrated to surpass a rigorous fault-tolerance threshold against general noise.
- The experimental system used traps with coherence times of approximately 1 s and trapping times ranging from several hours to 100 h.
B. Linear GST
Linear-inversion GST provides a reliable but low-accuracy initial estimate of the gate set, which seeds refinement by long-sequence GST.
- Linear-inversion GST provides a reliable but low-accuracy initial estimate that seeds subsequent long-sequence GST refinement.It is essentially simultaneous uncalibrated process and state tomography.
C. Analyzing long sequences in GST
Long-sequence GST combines iterative minimum-χ2 fitting with final maximum-likelihood estimation to improve accuracy and avoid local minima. Repeated sequences amplify gate errors, although stochastic decoherence limits useful sequence lengths.
- GST first fits short sequences, then successively adds longer sequences before performing a final maximum-likelihood fit using all data.The first stage uses minimum-χ2 estimation; the second is seeded from it and consistently avoids local minima.
- The iterative χ2 stage minimizes discrepancies between gate-set-predicted probabilities and observed frequencies for each experiment.Each experiment has predicted plus probability p_s, observed plus frequency f_s, and N samples.
- Long sequences amplify gate errors proportional to sequence length, reducing estimation error by a factor of L.This improves scaling from O(1/√N) to O(1/(L√N)) for suitable sequence lengths.
- Useful sequence lengths are bounded by stochastic decoherence, with the scaling breaking down for L ≥ 1/ϵ.In these experiments, ϵ ≤ 10^-4 and sequences reach L = 8192 ≈ 10^4.
- A hybrid min-χ2/MLE algorithm combines χ2’s numerical speed and stability with MLE’s statistical motivation and lack of bias.The χ2 estimate provides a reliable seed for the final MLE.
D. Selecting gate sequences for GST
GST uses fiducials and repeated germ sequences to make gate-set parameters observable and amplify their effects. Germ sets are selected through Jacobian sensitivity so that all gauge-invariant parameters are amplified.
- Each GST sequence sandwiches a repeated germ power between preparation and measurement fiducials.Six fiducials and 11 germs are used in the final GST runs.
- Fiducials provide informationally complete input states and measurements for probing the operation of interest.For single-qubit GST, six stabilizer-state fiducials are chosen as a convenient uniformly informationally complete set.
- Repeating a gate L times amplifies some deviations by L, but simple repetition can leave errors such as axis tilts unamplified.The tilt error in Gx cancels after four repetitions, requiring a more suitable germ such as GxGy.
- Germ selection computes Jacobians for candidate short sequences to identify parameter combinations each germ amplifies.The selection goal is high sensitivity at large L, using reversible-unitary assumptions to analyze the L → ∞ limit.
- A germ set is amplificationally complete when its Jacobian has right-singular rank equal to the number of physically accessible gate-set parameters.Jacobian singular vectors identify amplified parameter combinations, while singular values quantify amplification.
E. The GST gauge, and how to set it
Gate-set parameters include gauge freedoms that leave observable probabilities unchanged, complicating comparisons between representations. GST therefore gauge-optimizes the complete gate set against a target using a shared criterion.
- A gate set contains an initial state, measurement effect, and gates, but some representation parameters are physically unobservable gauge degrees of freedom.Gauge transformations can map distinct gate sets to identical experimental probabilities.
- Gauge freedom makes direct comparison difficult because fidelity, trace-norm distance, and diamond-norm distance are gauge-variant.The paper notes that few convenient gauge-invariant metrics are available.
- GST gauge-optimizes a gate set by choosing a transformation that minimizes a weighted Frobenius distance to a target gate set.The SPAM-to-gate weight ratio tunes the relative contributions of SPAM and logic-gate discrepancies.
- All reported metrics use one gauge optimized for the gate set as a whole, rather than separately optimizing each quantity.Separate optimization could give incompatible best gauges for different metrics.
F. Error bars
GST primarily uses Hessian-based likelihood-ratio confidence regions for uncertainty estimates, with bootstrapping as a cross-check and for RB quantities. The paper distinguishes parameter-wise intervals from joint confidence regions.
- Parametric bootstrapping generates datasets from the GST estimate with the same experiments and shot counts as the original data.Non-parametric bootstrapping instead resamples the experimental dataset with replacement.
- GST derives likelihood-ratio confidence regions from the Hessian of the log likelihood near the maximum-likelihood estimate.The Hessian defines a covariance tensor and an ellipsoidal 1 − α confidence region.
- 95% confidence intervals for scalar quantities are obtained by linearizing the quantity and projecting the Hessian onto non-gauge parameters.The intervals apply to quantities including fidelities, diamond norms, and gate matrix elements.
- The reported parameter-wise 95% intervals do not form a 95% joint confidence region for the entire gate set.With roughly 34 gauge-invariant parameters, the stated lower bound is 0.95^34 ≈ 17%.
- Bootstrapping agrees closely with likelihood-ratio error bars for gate elements and supplies uncertainty estimates for RB decay rates.RB decay rates are not amenable to likelihood-ratio regions because they are model-free.
VI. AUTHOR CONTRIBUTIONS
The authors divided responsibilities across theoretical development and code implementation, experiments, and manuscript-wide discussion and writing.
- RBK, JKG, EN, and KR contributed to the theoretical development of GST and its code implementation.
- JM, KF, and PM performed the experiments.
- All authors discussed the results and wrote the manuscript.