Source-linked AI summary
A Fault-Tolerant Honeycomb Memory
Craig Gidney, Michael Newman, Austin Fowler, Michael Broughton
TL;DR
The paper asks how robust the honeycomb quantum memory is and how many physical qubits it needs for trillion-operation reliability. It uses Monte Carlo simulations with a correlated minimum-weight perfect-matching decoder across circuit error models, comparing the honeycomb and rotated surface codes. The honeycomb threshold is 0.2%−0.3% in a standard circuit model, rises to 1.5%−2.0% with native two-body measurements, and reaches the teraquop regime with 600 qubits at 10^-3 physical error rate.
Problem
The paper evaluates whether a two-local honeycomb memory can provide robust logical-qubit protection and practical finite-size resource requirements despite locality-related error-correction costs.
Method
The authors use numerical simulations with a correlated minimum-weight perfect-matching decoder to estimate thresholds and teraquop qubit counts for the honeycomb and surface codes across error models.
Results
0.2%−0.3% honeycomb thresholds occur in the standard circuit model versus 0.5%−0.7% for the surface code, while native two-body measurements yield 1.5%−2.0% and a projected 600-qubit teraquop memory at 10^-3.
Takeaways & Limitations
The honeycomb code is exceptionally robust among two-local codes with entangling unitary circuits and among all codes considered with primitive two-body measurements.
Takeaways & Limitations
The honeycomb code does not achieve full effective distance in the EM3 model because correlated direct-measurement errors halve its effective distance.
Abstract
from arXiv · showhide
Recently, Hastings & Haah introduced a quantum memory defined on the honeycomb lattice. Remarkably, this honeycomb code assembles weight-six parity checks using only two-local measurements. The sparse connectivity and two-local measurements are desirable features for certain hardware, while the weight-six parity checks enable robust performance in the circuit model. In this work, we quantify the robustness of logical qubits preserved by the honeycomb code using a correlated minimum-weight perfect-matching decoder. Using Monte Carlo sampling, we estimate the honeycomb code's threshold in different error models, and project how efficiently it can reach the "teraquop regime" where trillions of quantum logical operations can be executed reliably. We perform the same estimates for the rotated surface code, and find a threshold of $0.2\%-0.3\%$ for the honeycomb code compared to a threshold of $0.5\%-0.7\%$ for the surface code in a controlled-not circuit model. In a circuit model with native two-body measurements, the honeycomb code achieves a threshold of $1.5\% < p <2.0\%$, where $p$ is the collective error rate of the two-body measurement gate - including both measurement and correlated data depolarization error processes. With such gates at a physical error rate of $10^{-3}$, we project that the honeycomb code can reach the teraquop regime with only $600$ physical qubits.
1 Introduction
The honeycomb code uses two-local measurements to assemble weight-six parity checks, addressing the absence of protected logical qubits in the static honeycomb subsystem code. This work numerically evaluates its thresholds and finite-size qubit costs against the surface code across error models.
- 1 Introduction: The honeycomb code assembles weight-six parity checks from two-local measurements on a sparse lattice, while its dynamic logical qubits change across measurement rounds.The construction uses three alternating measurement sub-rounds, and logical observables evolve as edge measurements are performed.
- 1 Introduction: The static subsystem code defined directly from the honeycomb constraints encodes no protected logical qubits, motivating the dynamic construction.Hastings and Haah’s construction instead encodes logical qubits into pairs of macroscopic anticommuting observables that change in time.
- 1.1 Summary of Results: 0.1%−0.3% thresholds make the honeycomb code approximately half as robust as the surface code under two-qubit unitary entangling gates.The comparison uses the same error model for both codes.
- 1.1 Summary of Results: At 10^-3 error rates, the honeycomb memory requires 5×−10× more qubits than the surface code in the considered baseline comparison, with a larger gap in the superconducting-inspired model.The superconducting-inspired gap is associated with assembling each honeycomb parity check from twelve measurements rather than two for the surface code.
- 1.1 Summary of Results: 1.5%−2.0% threshold performance emerges for the honeycomb code with native two-body measurements, compared with a previously reported 0.237% surface-code threshold.The prior surface-code result used a less accurate union-find decoder, so the thresholds are not directly comparable.
- 1.1 Summary of Results: 600 physical qubits are projected to reach the teraquop regime at a physical error rate of 10^-3 with native two-body measurements.Direct data-qubit measurements reduce noisy operations and eliminate measurement ancillas in the syndrome cycle.
2 Error Models
The study evaluates three circuit-level error models and compares thresholds, lambda factors, and teraquop qubit counts. These metrics characterize asymptotic suppression and finite-size resource requirements under different noisy gate architectures.
- 2 Error Models: Three noisy gate sets model standard depolarizing circuits, superconducting-inspired hardware, and primitive entangling measurements.The models are denoted SD6, SI1000, and EM3, respectively.
- 2 Error Models: SD6 is selected for comparison with quantum-error-correction literature, EM3 represents hardware with measurement as the primitive entangling operation, and SI1000 models superconducting-chip timing.The superconducting-inspired model assumes measurement and reset take about an order of magnitude longer than other gates.
- 2.1 Figures of merit: The threshold p_thr is the physical error rate below which some code size can achieve any target logical error rate.The required number of qubits may be potentially very large.
- 2.1 Figures of merit: The lambda factor Λ(p) measures error suppression obtained by increasing code distance by two and is approximately one at threshold.It additionally depends on the physical error rate p and is roughly distance-independent.
- 2.1 Figures of merit: The teraquop qubit count T(p) is the number of physical qubits needed for one trillion logical idle gates in expectation at physical error rate p.It combines error suppression with the total qubit cost of realizing the code.
3 Methods
The method combines Stim circuit simulation with a correlated minimum-weight perfect-matching decoder to evaluate honeycomb-code memory experiments. Stim generates detector error models and matching graphs, while Monte Carlo samples provide detection events for logical-observable decoding.
- Circuit generation: Stim generates circuit files with deterministic measurement-set detectors, logical-observable annotations, and annotated error channels for memory experiments.These circuit files describe the stabilizer circuit and the information needed to identify detection events and logical frame changes.
- Error-model construction: Stim converts each circuit into a detector error model by identifying violated detectors and flipped logical observables for every error mechanism.The resulting model represents each error through its detection-event symptoms and logical frame changes.
- Correlated decoding: The decoder converts detector error models into weighted matching graphs and reweights decomposed multi-symptom errors to capture correlations.Two- and one-event symptom sets become graph edges or boundary edges, while larger sets receive correlated-decoding reweighting rules.
- Sampling and evaluation: Stim samples detection events at high rate, and decoding succeeds when the inferred logical-observable outcome matches the initialized +1 eigenstate.This defines the logical-memory success criterion used in the simulations.
- Matching-graph structure: The honeycomb matching graph contains separate connected components for errors commuting or anticommuting with the preserved observable, with logical failures requiring at least four traversed edges.The graph shown for the 10-round 4x6-data-qubit SD6 circuit has degree at most 12 and uses green edges to mark errors flipping the annotated observable.
4 Results and Discussion
The honeycomb code has lower thresholds than the surface code in standard circuit models but reverses that comparison with native two-body measurements. Its hardware-efficient architecture can reach the teraquop regime with 600 qubits at a physical error rate of 10^-3, while correlated decoding improves below-threshold suppression.
- Thresholds: 0.2%-0.3%: The honeycomb threshold in SD6 is below the surface code’s 0.5%-0.7% threshold, while SI1000 gives 0.1%-0.15% versus 0.3%-0.5%.The relative reduction in SI1000 may reflect the proportionally noisier measurement channel and twelve-measurement honeycomb detectors.
- Thresholds: 1.5%-2%: The honeycomb code threshold in EM3 exceeds the 0.237% surface-code threshold reported for the same error model, despite decoder differences limiting direct comparability.The honeycomb advantage arises from direct two-body measurements, reduced circuit noise, and weight-six parity checks.
- Limitations: The EM3 effective distance is halved because correlated errors from two-body measurements overlap substantially with the observable.Both codes appear to achieve full effective distance in SD6 and SI1000, but the honeycomb code does not in EM3.
- Error suppression: Correlated decoding boosts below-threshold error suppression, even though it does not significantly increase the threshold.The lambda factor measures suppression when code distance increases by two; honeycomb distance changes require interpreting the plotted values as square roots of the four-step increase.
- Teraquop regime: 600 qubits: In EM3 at a physical error rate of 10^-3, the honeycomb code achieves a teraquop memory, whereas the surface code is more qubit-efficient in SD6 and SI1000.The teraquop count combines error suppression with the total qubits needed to realize the code.
- Teraquop regime: Teraquop qubit counts fall sharply near threshold as error rates decrease, but gains taper far below threshold, making lambda 10 more valuable than lambda 2 and lambda 50 less valuable than lambda 10.This reflects the shape of the projected teraquop curves rather than threshold alone.
5 Conclusions
The paper numerically tests honeycomb-code robustness against the surface code and introduces teraquop qubit count as a finite-size performance metric. It also identifies boundary-condition and layout choices that limit direct interpretation of the reported qubit estimate.
- Numerical tests find the honeycomb code exceptionally robust among two-local codes with entangling unitary circuits and among all codes with primitive two-body measurements.The surface code is also tested in two models to provide teraquop-regime baselines.
- The estimated teraquop qubit count is sensitive to boundary conditions, including whether the second logical qubit is counted.The work does not count the second logical qubit because its viability in a complete fault-tolerant computational system is uncertain.
- A rotated or sheared periodic layout could use 25% fewer physical qubits at the same code distance, but it was not simulated and may change the lambda factor.The authors therefore do not adjust the reported teraquop estimate using this layout.
6 Author Contributions
The authors divide the work across code implementation, simulation, decoding, surface-code comparison, and paper writing.
- Craig Gidney defined, verified, simulated, and plotted the honeycomb code circuits.
- Michael Newman reviewed the implementation and simulation results, generated surface-code comparison circuits, and handled most of the paper writing.
- Austin Fowler wrote the high-performance minimum-weight perfect-matching decoder and helped check simulation results.
- Michael Broughton helped scale the simulations over many machines.
A Detection Event Fractions
Detection event fractions quantify how often detectors fire across simulated circuits, code sizes, and noise parameters. The EM3 honeycomb model behaves differently from other models, while initialization and terminal measurement phases explain a small-distance uptick.
- The EM3 honeycomb error model has a significantly lower detection fraction than the other models.A tweaked EM3 model separates measurement-result flips from two-qubit data depolarization and produces detection fractions closer to other models.
- Detection fraction is the proportion of actual detection events among all potential detection events.
- At higher code distances, a slight initial uptick in detection-event fraction quickly stabilizes because nearby initialization and terminal measurement phases are less noisy.
B Converting Disjoint Error Mechanisms into Independent Error Mechanisms
The EM3 parity-measurement error mechanism uses disjoint cases, but Stim requires independent error mechanisms. The paper converts the model by choosing an independent-error probability that reproduces the original maximally mixing distribution.
- Stim-compatible error models require converting the EM3 disjoint error cases into independent error mechanisms.The conversion addresses the mismatch between the parity-measurement mechanism and Stim's independent-channel representation.
- The conversion uses n basis errors, with n = 5 for X1, Z1, X2, Z2, and flip.
- pmix is the probability of selecting one component error in a maximally mixing channel, while pind is the probability assigned independently to each error.The independent channels are chosen so their concatenation reproduces the maximally-mixing channel's error distribution.
C Additional Data
This section provides exhaustive supplementary plots, simulation data, and reproducibility resources for the evaluated honeycomb- and surface-code cases.
- The collected data and figure-generating Python code are available as ancillary files, including a CSV summarizing shots and errors for each experimental case.
- The simulations stopped at 100 million shots, 1000 errors, or clear above-threshold evidence, using roughly 10 CPU core years overall.
- Correlated-matching results are harder to reproduce because the decoder was an internal tool using hand-tuned heuristics.
- Exhaustive line-fit plots project logical error rates as code distance increases across error rates, gate sets, and decoding strategies.
- Exhaustive threshold plots cover both code types, multiple noisy gate sets, and standard or correlated decoding.
D Example Honeycomb Circuit
The example circuit implements a fault-tolerant honeycomb memory experiment, with annotated noise, detectors, observables, and repeated sub-rounds.
- The circuit fault-tolerantly initializes and measures the horizontal observable while preserving it through noise.
- The example uses 2×6 data qubits and 335 rounds, with the final round terminated before its second sub-round.
- Coordinate, noise, detector, and observable annotations are included, using the tweaked EM3 noise model from Figure 8.
- The circuit is provided as the ancillary file example_honeycomb_circuit.stim for reliable reuse.
- The repeated circuit alternates X, Y, and Z sub-rounds to compare stabilizer measurements and data measurements across rounds.