Source-linked AI summary

Suppressing quantum errors by scaling a surface code logical qubit

Rajeev Acharya, Igor Aleiner, Richard Allen, Trond I. Andersen, Markus Ansmann, Frank Arute, Kunal Arya, Abraham Asfaw, Juan Atalaya, Ryan Babbush, Dave Bacon, Joseph C. Bardin, Joao Basso, Andreas Bengtsson, Sergio Boixo, Gina Bortoli, Alexandre Bourassa, Jenna Bovaird, Leon Brill, Michael Broughton, Bob B. Buckley, David A. Buell, Tim Burger, Brian Burkett, Nicholas Bushnell, Yu Chen, Zijun Chen, Ben Chiaro, Josh Cogan, Roberto Collins, Paul Conner, William Courtney, Alexander L. Crook, Ben Curtin, Dripto M. Debroy, Alexander Del Toro Barba, Sean Demura, Andrew Dunsworth, Daniel Eppens, Catherine Erickson, Lara Faoro, Edward Farhi, Reza Fatemi, Leslie Flores Burgos, Ebrahim Forati, Austin G. Fowler, Brooks Foxen, William Giang, Craig Gidney, Dar Gilboa, Marissa Giustina, Alejandro Grajales Dau, Jonathan A. Gross, Steve Habegger, Michael C. Hamilton, Matthew P. Harrigan, Sean D. Harrington, Oscar Higgott, Jeremy Hilton, Markus Hoffmann, Sabrina Hong, Trent Huang, Ashley Huff, William J. Huggins, Lev B. Ioffe, Sergei V. Isakov, Justin Iveland, Evan Jeffrey, Zhang Jiang, Cody Jones, Pavol Juhas, Dvir Kafri, Kostyantyn Kechedzhi, Julian Kelly, Tanuj Khattar, Mostafa Khezri, Mária Kieferová, Seon Kim, Alexei Kitaev, Paul V. Klimov, Andrey R. Klots, Alexander N. Korotkov, Fedor Kostritsa, John Mark Kreikebaum, David Landhuis, Pavel Laptev, Kim-Ming Lau, Lily Laws, Joonho Lee, Kenny Lee, Brian J. Lester, Alexander Lill, Wayne Liu, Aditya Locharla, Erik Lucero, Fionn D. Malone, Jeffrey Marshall, Orion Martin, Jarrod R. McClean, Trevor Mccourt, Matt McEwen, Anthony Megrant, Bernardo Meurer Costa, Xiao Mi, Kevin C. Miao, Masoud Mohseni, Shirin Montazeri, Alexis Morvan, Emily Mount, Wojciech Mruczkiewicz, Ofer Naaman, Matthew Neeley, Charles Neill, Ani Nersisyan, Hartmut Neven, Michael Newman, Jiun How Ng, Anthony Nguyen, Murray Nguyen, Murphy Yuezhen Niu, Thomas E. O'Brien, Alex Opremcak, John Platt, Andre Petukhov, Rebecca Potter, Leonid P. Pryadko, Chris Quintana, Pedram Roushan, Nicholas C. Rubin, Negar Saei, Daniel Sank, Kannan Sankaragomathi, Kevin J. Satzinger, Henry F. Schurkus, Christopher Schuster, Michael J. Shearn, Aaron Shorter, Vladimir Shvarts, Jindra Skruzny, Vadim Smelyanskiy, W. Clarke Smith, George Sterling, Doug Strain, Marco Szalay, Alfredo Torres, Guifre Vidal, Benjamin Villalonga, Catherine Vollgraff Heidweiller, Theodore White, Cheng Xing, Z. Jamie Yao, Ping Yeh, Juhwan Yoo, Grayson Young, Adam Zalcman, Yaxing Zhang, Ningfeng Zhu

arXiv:2207.06431v2quant-ph

TL;DR

Practical quantum computing needs much lower error rates than physical qubits currently provide, motivating scalable quantum error correction. This paper measures logical performance across surface-code sizes and finds modest improvement for distance-5 over distance-3, while identifying high-energy events and correlated errors as barriers to further scaling.

  • Problem

    The paper asks whether increasing error-correcting code size reduces logical error rates in a real device despite the additional physical error sources.

  • Method

    The study compares distance-5 and averaged distance-3 surface-code logical qubits on superconducting hardware, models their physical noise, and tests a distance-25 repetition code.

  • Results

    2.914% ± 0.016% logical error per cycle for distance-5 compares with 3.028 ± 0.023% for averaged distance-3 codes, while distance-25 repetition coding reaches (1.7 ± 0.3) × 10^-6 per cycle without post-selection.

  • Takeaways & Limitations

    The results demonstrate that quantum error correction begins improving logical performance with increasing qubit number, while highly correlated errors must be mitigated for scalable correction.

Abstract

from arXiv · show

Practical quantum computing will require error rates that are well below what is achievable with physical qubits. Quantum error correction offers a path to algorithmically-relevant error rates by encoding logical qubits within many physical qubits, where increasing the number of physical qubits enhances protection against physical errors. However, introducing more qubits also increases the number of error sources, so the density of errors must be sufficiently low in order for logical performance to improve with increasing code size. Here, we report the measurement of logical qubit performance scaling across multiple code sizes, and demonstrate that our system of superconducting qubits has sufficient performance to overcome the additional errors from increasing qubit number. We find our distance-5 surface code logical qubit modestly outperforms an ensemble of distance-3 logical qubits on average, both in terms of logical error probability over 25 cycles and logical error per cycle ($2.914\%\pm 0.016\%$ compared to $3.028\%\pm 0.023\%$). To investigate damaging, low-probability error sources, we run a distance-25 repetition code and observe a $1.7\times10^{-6}$ logical error per round floor set by a single high-energy event ($1.6\times10^{-7}$ when excluding this event). We are able to accurately model our experiment, and from this model we can extract error budgets that highlight the biggest challenges for future systems. These results mark the first experimental demonstration where quantum error correction begins to improve performance with increasing qubit number, illuminating the path to reaching the logical error rates required for computation.

I. INTRODUCTION

Quantum applications require billions of operations, but physical gate error rates remain too high for such circuits. Quantum error correction can suppress operational errors, yet it requires substantial physical-qubit overhead and must improve as code size grows.

  • Billions of quantum operations are often required for applications including factoring, optimization, machine learning, simulation, and quantum chemistry.
  • 10^-3 per gate is typical for state-of-the-art quantum processors, which is too high for executing such large circuits.
  • Quantum error correction can exponentially suppress operational error rates at the expense of physical-qubit overhead.
  • 72-qubit superconducting hardware supports a 49-qubit distance-5 surface code that narrowly outperforms its average 17-qubit distance-3 subset codes.

II. SURFACE CODES WITH SUPERCONDUCTING QUBITS

Surface codes encode a logical qubit non-locally across data qubits and detect errors through repeated stabiliser measurements. The experiment implements distance-5 and distance-3 codes on a superconducting device to compare logical performance as code size increases.

  • Surface codes encode a logical qubit in a d × d square of entangled data qubits using anti-commuting observables X_L and Z_L.
  • Adjacent data-qubit parities are periodically measured with d^2−1 measure qubits, whose stabilisers commute with the logical observables and one another.
  • A decoder uses stabiliser-measurement histories to infer physical errors and determine their overall effect on the logical qubit.
  • 49 qubits implement the distance-5 code, comprising 25 data qubits and 24 measure qubits.
  • The distance-5 code is compared with four minimally overlapping distance-3 subgrids, each containing 9 data qubits and 8 measure qubits.
  • Each error-correction cycle sequences CZ and Hadamard gates to extract X and Z stabilisers simultaneously, followed by measure-qubit measurement and reset.

III. ERROR DETECTORS

Detection events identify inconsistencies in repeated stabiliser parities and reveal how physical errors are distributed across space and time. Experiment and simulation show comparable detector rates across code distances, while leakage and related interactions produce a time-dependent rise.

  • A detection event occurs when a stabiliser parity differs from the preceding cycle beyond known circuit-induced flips.
  • 25-cycle experiments over 50,000 instances measure detection probabilities for distance-5 and four distance-3 codes.
  • 0.185 ± 0.018 is the distance-5 weight-4 detection probability, compared with 0.175±0.017 averaged across distance-3 codes.
  • 0.119 ± 0.012 is the distance-5 weight-2 detection probability, compared with 0.115 ± 0.008 averaged across distance-3 codes.
  • Detection probabilities rise 12% for distance-5 and 8% for distance-3 over 25 cycles, with a characteristic risetime of roughly 5 cycles.
  • The Pauli+ simulation predicts distance-5 averages of 0.180 ± 0.013 for weight-4 and 0.116 ± 0.011 for weight-2 stabilisers, with a 7% rise over 25 rounds.

IV. UNDERSTANDING ERRORS THROUGH CORRELATIONS

Correlations between detection events provide finer-grained information about physical error mechanisms than individual detection rates. Experimental correlations exceed those from a Pauli-only model, while a richer model incorporating leakage and stray interactions better matches the data.

  • Pairwise detection correlations classify measurement and reset errors as timelike, data-idling errors as spacelike, and some gate errors as diagonal pairs.
  • Y errors generate more complex detection-event clusters by producing events associated with both X and Z errors.
  • The Pauli-only simulation systematically underpredicts experimental correlation probabilities, whereas Pauli+ is closer and predicts unexpected pairs.
  • Leakage and stray interactions can create distantly separated detection events that a decoder might misinterpret as multiple independent errors.

V. DECODING AND LOGICAL ERROR PROBABILITIES

The study uses error-hypergraph-based belief-matching and tensor-network decoders to evaluate logical performance, finding that distance-5 modestly improves over averaged distance-3 codes as physical performance improves.

  • Decoding: Belief-matching updates physical-error probabilities using nearby detection events before applying matching, while tensor-network decoding approximates maximum likelihood over logical-error classes.The tensor-network decoder is slower but more accurate given the error-hypergraph priors.
  • Logical error probabilities: Over 25 cycles, distance-5 produces lower logical error probabilities than the average of four distance-3 subset codes.Logical errors are averaged across X and Z bases without post-selecting leakage or high-energy events.
  • Scaling progression: The larger code improved about twice as fast as the smaller code during physical-error-rate improvements, eventually overtaking distance-3 performance.The progression compares distance-5 with individual and averaged distance-3 codes across system improvements.
  • Error budget: CZ errors and data-qubit decoherence during measurement and reset are dominant contributors to the modeled logical-error budget.The budget varies physical error rates in simulations to estimate each mechanism’s contribution to 1/Λ.

VI. ALGORITHMICALLY-RELEVANT ERROR RATES WITH REPETITION CODES

The paper tests much lower logical error rates with a bit-flip repetition code and identifies high-energy impacts as a damaging correlated-error source. A distance-25 code reaches the 10^-6 regime, with lower error after excluding one such event.

  • Repetition-code test: The bit-flip repetition code corrects only bit-flip errors, making it unsuitable for quantum algorithms but capable of much lower logical error probabilities.Its restricted correction capability enables the low-error-rate test.
  • Repetition-code results: 1.7 ± 0.3 × 10^-6 logical error per cycle is achieved without post-selection using a distance-25 repetition code.The code is decoded with minimum-weight perfect matching.
  • Repetition-code results: After removing trials near one high-energy event, logical error per cycle decreases to 1.6 ± 0.8 × 10^-7.The event affected 0.15% of trials and temporarily produced widespread correlated errors.
  • Implications: The results identify mitigation of highly correlated errors such as cosmic-ray impacts as an important requirement for scalable quantum error correction.High-energy events can be identified through spikes in detection-event counts.

VII. TOWARDS LARGE-SCALE QUANTUM ERROR CORRECTION

Simulations place the experiment in a finite-size crossover regime where larger codes initially suppress logical errors before a turnaround. Reaching practical logical error rates will require substantially better components and experimental validation at larger scales.

  • Scaling regimes: At high physical error rates, increasing code size worsens logical error, whereas below threshold larger codes improve it.The simulations examine surface-code distances from 3 to 25 while scaling the measured physical error model.
  • Scaling regimes: The experiment lies in a crossover regime where increasing system size initially suppresses logical error before later increasing it.This behavior reflects finite-size effects in the current operating regime.
  • Future requirements: Component performance must improve by at least 20% to move below threshold and improve significantly further for practical scaling.This estimate is based on the error budget and simulations.
  • Future requirements: The projections rely on simplified models and require experimental validation with larger code sizes and longer durations.The paper presents this work as an initial step toward the desired logical performance.

VIII. AUTHOR CONTRIBUTIONS

The Google Quantum AI team conceived and designed the experiment, while theory and experimental teams developed the analysis, modeling, system, calibration, and data collection.

  • Author contributions: The Google Quantum AI team conceived and designed the experiment.
  • Author contributions: Google Quantum AI developed analysis, modeling, and metrological tools, built the system, performed calibrations, and collected data.Modeling was conducted jointly with outside collaborators.
  • Author contributions: All authors wrote and revised the manuscript and Supplementary Information.

X. DATA AVAILABILITY

The data supporting the paper’s plots and other findings are available upon reasonable request or through the listed Zenodo record. The document is attributed to Google Quantum AI and dated July 21, 2022.

  • Data supporting the paper’s plots and other findings are available upon reasonable request or through Zenodo record 10.5281/zenodo.6804040.
  • The paper is attributed to Google Quantum AI.
  • The document is dated July 21, 2022.

G. Surface code experimental details

The experiment uses calibrated surface-code circuits, randomized initial states, interleaved datasets, and decoder validation to compare code sizes while controlling drift and modeling uncertainty.

  • Circuit calibration: The surface-code circuit uses virtual Zα corrections before Hadamard gates to minimize detection probabilities from effective single-qubit phase shifts.Corrections are empirically optimized in compatible groups and typically satisfy α < 0.1.
  • Grid selection: Distance-3 grids were selected to minimize overlap, but their shared center qubit remains essential for a fair comparison with distance-5.Anomalous center-qubit phase errors can dramatically affect distance-3 performance and create an outlier-driven crossover.
  • Data acquisition: Each dataset contains 50,000 surface-code runs, and acquisition order is shuffled across cycle counts, bases, and codes.A frequency-update Ramsey calibration is performed after every five cycles.
  • State preparation: Random initial bitstrings prevent first-round measurement bias and artificial reductions in early-round error rates during code warmup.Ten bitstrings are used for the main data, with five for each logical value.
  • Decoder validation: Belief-matching remains robust when its prior is computed from earlier device data, with distance-5 logical error per cycle changing from 3.056% to 3.059%.The corresponding distance-3 change is from 3.118% to 3.129%.
  • Model validation: An extended fluctuating-error model produces nearly identical distance-3 versus distance-5 separation, supporting the robustness of the reported fit.A quantum-trajectories simulation independently models noise using explicit quantum channels rather than only Pauli channels.

C. Comparison of Pauli+ and quantum simulations

Pauli+ and quantum simulations generally agree on logical error and typical detection-event statistics, while the Pauli+ model captures the main observed correlations. The simulations use experimentally grounded component errors and sensitivity analysis to characterize logical-error behavior.

  • C. Comparison of Pauli+ and quantum simulations: 200,000 samples per 25-round simulation took about 160 seconds for Pauli+ versus 11 hours for the quantum simulation.The comparison used distance-3 Z- and X-basis West-device experiments.
  • C. Comparison of Pauli+ and quantum simulations: Good agreement appears between the simulations for logical error and typical detection-event statistics.The agreement is reported for the distance-3 West, X- and Z-basis comparisons.
  • C. Comparison of Pauli+ and quantum simulations: Stabilizer-state projection can effectively twirl quantum noise, making the approximation exact for isolated single-qubit errors without leakage.The explanation also applies to errors in non-neighboring circuit locations and to exactly modeled Pauli channels such as dephasing.
  • C. Comparison of Pauli+ and quantum simulations: Unitary crosstalk rotations may add coherently in quantum Monte Carlo simulations, complicating the approximation’s behavior.Earlier work reported overprediction under some decoherence and unitary-error conditions, whereas the present comparisons show generally good agreement.
  • C. Comparison of Pauli+ and quantum simulations: Logical-error sensitivity is evaluated by varying one component error probability and measuring the resulting change in logical error per cycle.The sensitivity is nonlinear overall, so coefficients depend on the experimental operation point.

A. Experimental operation point

The experimental operation point is built from independently measured component-error probabilities and modeled leakage, heating, and crosstalk processes. Crosstalk contributes roughly 15% of the observed total CZ error.

  • A. Experimental operation point: Component-error probabilities are independently measured for single-qubit gates, CZ gates, data-qubit idles, readout, and reset.Single-qubit errors use randomized benchmarking, CZ errors use cross-entropy benchmarking, and idle errors use interleaved randomized benchmarking.
  • A. Experimental operation point: Leakage is modeled as arising during CZ operations and from nonequilibrium heating into non-computational states.The nominal CZ leakage parameter uses p_t = 8 × 10^-4, while heating is characterized using Γ_1→2 = 1/(700 µs).
  • A. Experimental operation point: Crosstalk is modeled separately for simultaneous gates and idles, producing an averaged Pauli error probability for CZ gates.The crosstalk contribution is subtracted from total CZ error to obtain the depolarizing CZ component.
  • A. Experimental operation point: 15% of the observed total CZ error is attributed to the modeled crosstalk contribution.This is the reported average contribution across CZ gates.

B. Logical error rate sensitivities for d = 3, 5 surface codes

Sensitivity analysis quantifies how d = 3 and d = 5 logical error per cycle responds to component, leakage, and crosstalk errors. The d = 5 code is more sensitive to component errors, with sensitivities obtained from linearized simulation responses.

  • B. Logical error rate sensitivities for d = 3, 5 surface codes: d = 5 surface code is more sensitive to component errors than d = 3 surface code.The comparison includes sensitivities to leakage and crosstalk as well as the main circuit components.
  • B. Logical error rate sensitivities for d = 3, 5 surface codes: The logical error per cycle is nonlinear in component-error probabilities, so sensitivity coefficients depend on the chosen operation point.The reported coefficients are evaluated at the experimental operation point.
  • B. Logical error rate sensitivities for d = 3, 5 surface codes: Sensitivity coefficients are extracted from the slope of the change in logical error per cycle versus a component-error variation.For single-qubit gates, the variation is applied uniformly across the relevant gates at the experimental operation point.
  • B. Logical error rate sensitivities for d = 3, 5 surface codes: Heating sensitivity varies Γ_1→2, whereas crosstalk sensitivity scales all crosstalk unitaries through a parameter α.These variations are mapped to the error-probability perturbation used in the sensitivity formula.

A. Experiment vs. Pauli+ correlation matrices pij

The p_ij correlation analysis shows that Pauli+ simulation captures most detection-event correlations, while stronger experimental correlations reveal localized leakage discrepancies. Time-averaged edge probabilities provide component-specific diagnostics across the code.

  • A. Experiment vs. Pauli+ correlation matrices p_ij: Pauli+ simulation captures most observed detection-event correlations, leaving relatively faint residual differences from experiment.The largest residuals are associated with stronger leakage in data qubits 2_4 and 2_6.
  • A. Experiment vs. Pauli+ correlation matrices p_ij: Leakage accumulation in data qubits 2_4 and 2_6 creates local spatial and roughly T_1/2 ≈ 10 µs nonlocal temporal correlations.These correlations appear as broad residual regions near measure qubits neighboring the leaky data qubits.
  • A. Experiment vs. Pauli+ correlation matrices p_ij: Timelike edges primarily diagnose readout/reset errors, while spacelike and spacetimelike edges diagnose idle data-qubit, CZ, leakage, and crosstalk effects.The edge classes separate correlations by spatial and temporal separation between detection events.
  • A. Experiment vs. Pauli+ correlation matrices p_ij: The mean T-edge discrepancy is relatively small, with the largest mismatches near data qubits exhibiting excess leakage.Time-averaged SX and SZ edge probabilities also show relatively small experiment–simulation differences.

3. Spacetimelike edges

The analysis distinguishes expected spacetimelike correlations from unexpected ST’-edges, using experiment–simulation discrepancies to diagnose leakage and construct error budgets.

  • ST-edges are expected from CZ Pauli errors, whereas ST’-edges are unexpected correlations mainly attributed to leakage accumulation in data qubits.
  • The similar excess ST-edge and ST’-edge probabilities indicate that leakage contributes substantially to the experimentally observed ST-edge excess.
  • 0.6 × 10−2 is the experimental mean probability of unexpected ST’-edges, compared with 0.35 × 10−2 in Pauli+ simulation.
  • The largest experiment–simulation discrepancies occur for ST’-edges involving measure qubits spatially close to data qubits 2_4 and 2_6, where experiment shows more leakage accumulation.
  • The generalized pij correlation method characterizes higher-order detection-event clusters and supplies cluster probabilities for decoder weighting and error diagnostics.
Loading 2207.06431v2…