Source-linked AI summary

Scalable Neural Decoders for Practical Fault-Tolerant Quantum Computation

Andi Gu, J. Pablo Bonilla Ataides, Mikhail D. Lukin, Susanne F. Yelin

arXiv:2604.08358v1quant-phcs.AIcs.LG

TL;DR

Quantum error correction requires decoders that can interpret spacetime syndromes accurately and efficiently, but existing quantum decoders miss important error-suppression regimes. This paper develops a geometry-aware convolutional decoder and shows near-optimal suppression across code families, with implications for decoder–architecture co-design and practical fault tolerance.

  • Problem

    Quantum decoders must infer logical-error classes from spacetime syndromes, while existing decoders are not accurate enough to expose the full error-suppression behavior of quantum codes.

  • Method

    The decoder uses learned local message-passing updates with locality, translation equivariance, and anisotropy, implemented through successive layers that coarse-grain errors across scales.

  • Results

    The decoder achieves near-optimal error suppression on surface and quantum LDPC codes, including ∼17× lower logical error rates than Relay on the Gross code and Λ ≈ 8.4 on surface codes.

  • Takeaways & Limitations

    Code distance alone is insufficient for resource estimation, and decoder expressive power should be treated as part of fault-tolerant architecture co-design.

  • Takeaways & Limitations

    Local attention offers lower theoretical computational cost but currently has limited optimized CUDA support, while GPU latency estimates can poorly represent dedicated-hardware performance.

Abstract

from arXiv · show

Quantum error correction (QEC) is essential for scalable quantum computing. However, it requires classical decoders that are fast and accurate enough to keep pace with quantum hardware. While quantum low-density parity-check codes have recently emerged as a promising route to efficient fault tolerance, current decoding algorithms do not allow one to realize the full potential of these codes in practical settings. Here, we introduce a convolutional neural network decoder that exploits the geometric structure of QEC codes, and use it to probe a novel "waterfall" regime of error suppression, demonstrating that the logical error rates required for large-scale fault-tolerant algorithms are attainable with modest code sizes at current physical error rates, and with latencies within the real-time budgets of several leading hardware platforms. For example, for the $[144, 12, 12]$ Gross code, the decoder achieves logical error rates up to $\sim 17$x below existing decoders - reaching logical error rates $\sim 10^{-10}$ at physical error $p=0.1\%$ - with 3-5 orders of magnitude higher throughput. This decoder also produces well-calibrated confidence estimates that can significantly reduce the time overhead of repeat-until-success protocols. Taken together, these results suggest that the space-time costs associated with fault-tolerant quantum computation may be significantly lower than previously anticipated.

Structure-aware neural decoding

Cascade is designed around the regular geometric and spatiotemporal structure of surface and bivariate bicycle codes, replacing fixed message-passing rules with learned convolutional operations.

  • Surface and bivariate bicycle codes arrange stabilizers in regular spatial lattices that extend into translation-regular spatiotemporal structures under repeated measurements.
  • Cascade exploits locality, translation equivariance, and anisotropy to process spatially localized syndrome patterns across multiple scales.
  • Belief propagation uses fixed local update rules that can converge to incorrect solutions on degenerate quantum error patterns.
  • Cascade replaces fixed rules with learned convolutional operations that preserve the decoder’s structural priors.

Error suppression, speed and accuracy

Cascade exposes a steep waterfall regime in logical-error suppression across BB and surface codes, while combining strong accuracy with practical throughput and latency. The results indicate that decoder performance and code structure can substantially improve fault-tolerance resource estimates beyond distance-only scaling.

  • Cascade’s waterfall regime arises when numerous higher-weight failure modes dominate before rare minimum-weight modes produce a distance-limited floor.
  • ∼4000× lower logical error than BP+OSD and ∼17× lower than Relay are achieved on the Gross code at p = 0.1%.
  • PL ∼10−10 per logical qubit per cycle is achieved for J288,12,18K at p = 0.2%.
  • Λ ≈8.4 for Cascade exceeds MWPM at Λ ≈5.0 and correlated MWPM at Λ ≈7.8, approaching Tesseract at Λ ≈9.1.
  • A target logical error rate of ∼10−9 requires d = 15 for Cascade versus d = 19 for MWPM, corresponding to a ∼40% physical-qubit reduction.
  • Single-shot latency is ∼40 µs per cycle, while batched inference provides 3,000–100,000× higher throughput than existing single-threaded CPU decoders.

Robustness and generalization

Cascade generalizes across physical error rates while its confidence estimates remain calibrated, enabling effective post-selection with substantially fewer retries. Increasing model capacity is key to accessing the steep waterfall regime under circuit-level noise.

  • Capacity dependence: Capacity increases the suppression exponent to m ≈8 for H ≳64, while insufficient capacity prevents decoders from accessing the waterfall regime under circuit-level noise.Under circuit-level noise, higher-weight failure modes drive the waterfall because minimum-weight spacetime failures are relatively rare.
  • Confidence estimates: Calibration persists across physical error rates far below training, allowing confidence-based post-selection to reduce logical errors by discarding low-confidence predictions.The confidence estimates remain meaningful across the evaluation range, despite changes in the posterior error probability.
  • Confidence estimates: Up to two orders-of-magnitude reduction in PL is obtained with only a 0.5% per-cycle discard rate.The tradeoff is lower acceptance, making the result relevant to confidence-aware post-selection.

Discussion and outlook

The results argue for code- and decoder-aware resource estimation and for treating decoder design as part of fault-tolerant architecture co-design. Cascade’s geometric inductive bias extends across surface and bivariate bicycle codes, while current hardware error rates lie in the regime where the waterfall effect may be useful.

  • Resource estimation: Distance alone is insufficient for resource estimation because low-weight failure sparsity, logical-operator weight distributions, and decoder accuracy vary across codes with identical parameters.Code-specific models that capture these structural properties yield more favorable resource estimates than standard distance-based formulas.
  • Architecture co-design: Decoder capacity directly determines how much of a code’s error-correcting capability is realized in practice.Below a sharp expressive-capacity threshold, decoders miss complex error patterns; above it, near-optimal performance emerges.
  • General framework: The same geometric architecture achieves near-optimal accuracy on surface and bivariate bicycle codes despite their different connectivity and encoding rates.The framework uses locality, translation equivariance, and anisotropy rather than being tailored to one code family.
  • Hardware outlook: ∼0.1% entangling error rates on trapped-ion, superconducting, and neutral-atom systems place current hardware in the regime where the waterfall effect may enable useful logical error rates.The passage connects present hardware error rates with the paper’s identified waterfall regime.

METHODS

The paper develops a decoder framework for geometrically structured surface and bivariate bicycle codes, using learned convolutional message passing to exploit translation symmetry, locality, and anisotropy. Its architecture scales depth with code distance so that local operations can resolve error structure across the code.

  • Decoding problem: The decoding task infers whether a logical error occurred from repeated syndrome measurements and detection events.The reported logical error rate PL is the probability that the correction combined with actual errors produces a logical bit flip.
  • Geometric structure: Surface and bivariate bicycle codes arrange stabilizers in regular spatial structures that support translation-equivariant decoder design.Surface codes use a 2D grid, while bivariate bicycle codes use a torus; both permit weight-sharing across translated neighborhoods.
  • Decoder design: Cascade encodes binary detection events at syndrome locations and processes them with learned convolutional operations that preserve locality, translation equivariance, and anisotropy.Relative geometric offsets provide direction-specific message rules, unlike standard graph neural networks that treat neighbors symmetrically.
  • Decoder design: Locality is implemented as hierarchical coarse-graining: successive layers resolve errors at progressively larger characteristic scales.This motivates using approximately L ∼ d layers so the receptive field spans the distance-dependent correlations relevant to decoding.
  • Architectural comparison: The architectural gain is decomposed into locality first, followed by translation equivariance and anisotropy when moving from local attention to convolution.Restricting attention to local neighborhoods improves accuracy and compute efficiency; convolution adds identical direction-specific rules at every position.

SUPPLEMENTARY INFORMATION: TOWARD REAL-TIME DECODING ON DEDICATED HARDWARE

The supplementary analysis evaluates dedicated-hardware implementations of Cascade, covering quantization, architectural cost, latency, and practical deployment constraints. Convolution offers strong accuracy and regular computation, while depthwise variants and FP8 reduce hardware cost and approach real-time budgets on several platforms.

  • Quantization and inference optimization: FP8 inference preserves logical error rates indistinguishable from full precision across tested physical error rates without quantization-aware retraining.Batch-normalization parameters are folded into convolution weights, leaving a pure convolutional inference graph.
  • Quantization and inference optimization: ∼60× reduction in multiplier area results from combining depthwise convolution’s ∼4× reduction with FP8’s ∼16× reduction versus standard FP32 convolution.Further INT4 reduction is identified as future work.
  • Computational cost and roofline latency: Convolution achieves the best accuracy in the ablation, whereas local attention uses fewer MACs but has degraded accuracy at fixed model size.The comparison decomposes each block into two pointwise projections and one spatial operation.
  • Computational cost and roofline latency: ∼4× net speedup comes from depthwise factorization at representative parameters, because unchanged pointwise projections become the bottleneck after reducing spatial cost.The spatial term accounts for ∼77% of standard-convolution MACs in this example.
  • Computational cost and roofline latency: Even the largest models fit within the ∼1 ms decoding budgets of trapped-ion and neutral-atom platforms, while standard H = 256 convolutions miss the ∼1 µs superconducting target.Depthwise convolutions with H ≤128 approach the superconducting target; these are roofline estimates assuming 100% utilization.
  • Dedicated-hardware deployment: Regular feed-forward computation enables spatial dataflow mapping onto FPGA or ASIC fabric, with layers assigned to dedicated hardware blocks.The implementation avoids dynamic scheduling and control flow; on GPUs, unbatched latency can instead be dominated by launch overhead and memory transfers.
  • Dedicated-hardware deployment: Fixed-depth convolutional inference provides deterministic latency, unlike iterative or data-dependent decoders whose execution paths can vary.This property is relevant to real-time QEC systems that are sensitive to worst-case latency.
Loading 2604.08358v1…