Source-linked AI summary

The Pinnacle Architecture: Reducing the cost of breaking RSA-2048 to 100 000 physical qubits using quantum LDPC codes

Paul Webster, Lucas Berent, Omprakash Chandra, Evan T. Hockings, Nouédyn Baspin, Felix Thomsen, Samuel C. Smith, Lawrence Z. Cohen

arXiv:2602.11457v2quant-ph

TL;DR

Fault-tolerant quantum architectures are needed to overcome the precision requirements and noise affecting engineered quantum systems. The paper introduces the Pinnacle Architecture using QLDPC codes and reports RSA-2048 factoring with fewer than one hundred thousand physical qubits under stated hardware assumptions, alongside broader utility-scale applications.

  • Problem

    Quantum computing's potential for efficient solutions to currently intractable problems depends on fault-tolerant architectures because engineered systems experience significant noise and require high precision.

  • Method

    The Pinnacle Architecture uses bridged QLDPC code blocks, modular generalised-surgery gadgets, a magic engine, Clifford frame cleaning, and flexible parallel operation for universal quantum computation.

  • Results

    Fewer than one hundred thousand physical qubits suffice to factor 2048-bit RSA integers at p = 10−3, with 1 µs code cycles and a 10 µs reaction time.

  • Takeaways & Limitations

    The architecture opens the possibility of utility-scale quantum computing on one hundred thousand physical qubit devices and may hasten practical quantum computing.

  • Takeaways & Limitations

    The reported architecture uses generalised bicycle QLDPC codes, while higher-rate QLDPC codes and further component optimisation remain possible avenues for additional reductions.

Abstract

from arXiv · show

The realisation of utility-scale quantum computing inextricably depends on the design of practical, low-overhead fault-tolerant architectures. We introduce the Pinnacle Architecture, which uses quantum low-density parity check (QLDPC) codes to allow for universal, fault-tolerant quantum computation with a spacetime overhead significantly smaller than that of any competing architecture. With this architecture, we show that 2048-bit RSA integers can be factored with fewer than one hundred thousand physical qubits, given a physical error rate of $10^{-3}$, code cycle time of $1$ microsecond and a reaction time of $10$ microseconds. We thereby demonstrate the feasibility of utility-scale quantum computing with an order of magnitude fewer physical qubits than has previously been believed necessary.

I. INTRODUCTION

The Pinnacle Architecture targets the high overhead of fault-tolerant quantum computing by combining QLDPC-based modular components, efficient logical measurements, and parallelism. Its benchmark estimates indicate order-of-magnitude resource reductions for RSA factoring and Fermi-Hubbard ground-state estimation across hardware regimes.

  • Motivation: Surface-code architectures may require at least one million physical qubits because hundreds or thousands encode each low-failure-rate logical qubit.Scaling quantum hardware to this size poses formidable challenges.
  • Architecture: The Pinnacle Architecture uses bridged QLDPC code blocks, generalised-surgery gadgets, magic engines, and Clifford frame cleaning to reduce spacetime overhead.Its magic engine supports distillation and injection in a single code block, while Clifford frame cleaning enables parallel operations across processing units.
  • Benchmark results: Fewer than one hundred thousand physical qubits suffice for RSA-2048 factoring at p = 10^-3, a 1 µs code cycle time, and a 10 µs reaction time.The estimate compares with close to one million physical qubits in the previous best result; parallelisation also supports alternative space-time trade-offs.
  • Architecture: Modular processing units support limited connectivity, arbitrary logical Pauli measurements, universal computation, and parallel read-only access to shared quantum memory.The architecture requires interactions only on the scale of a processing block, constant in the number of logical qubits.
  • Benchmark results: 58 thousand and 20 thousand physical qubits suffice for Fermi-Hubbard instances at L = 16 with p = 10^-3 and p = 10^-4, respectively.These compare with 940 thousand and 200 thousand physical qubits in Ref. [13], while maintaining runtimes from minutes to days depending on code cycle time.
  • Benchmark results: The architecture's example configurations trade runtime against physical-qubit count by changing processing-unit parallelisation, code cycle time, and physical error rate.One example factors RSA-2048 in one month with approximately one hundred thousand physical qubits, while another uses 81 processing units and approximately one million physical qubits for a three-month regime.

III. BACKGROUND

The paper frames fault-tolerant quantum architectures as necessary for useful quantum computing, then reviews QLDPC-based building blocks, logical operations, compilation, and runtime timescales.

  • A. Code Blocks: A code block performs repeated syndrome extraction, while a logical cycle combines Θ(d) code cycles to obtain a reliable error syndrome.
  • QLDPC codes bound check weights and qubit degrees independently of code distance, enabling constant-depth syndrome extraction circuits.
  • B. Processing Blocks: Measurement gadgets and bridges convert QLDPC code blocks into processing blocks supporting arbitrary logical Pauli measurements during error correction.
  • Pauli-based computation compiles Clifford+T circuits into logical measurements and gives a time cost that scales with T count.
  • D. Relevant Timescales: Runtime estimates use code-cycle, logical-cycle, and reaction times; the architecture assumes tr = 10tc and remains non-reaction-limited when dt ≥ 10.

IV. THE PINNACLE ARCHITECTURE

The Pinnacle Architecture combines QLDPC processing units, continuously operating magic engines, and optional memory into a modular, parallelisable system for fault-tolerant computation.

  • The architecture is built from processing units, magic engines, and optional memory modules based on QLDPC codes.
  • 1. Processing Units: Each processing unit uses bridged QLDPC processing blocks to support arbitrary logical Pauli measurements across its encoded logical qubits.
  • 2. Magic Engines: Magic engines simultaneously produce and consume magic states, providing continuous throughput to an associated processing unit.
  • 2. Magic Engines: Each magic engine produces one encoded |T̄⟩ state per logical cycle while injecting the previous cycle’s state into its processing unit in parallel.
  • 2. Magic Engines: Magic-state distillation uses one logical sector, while the other sector supports parallel consumption through joint Pauli measurements with the processing unit.
  • 2. Magic Engines: Distillation rejection increases the expected logical cycles per T gate from 1 to α = (1 − pr)^−1, and may leave the processing unit idle.
  • 3. Memory: Optional memory stores logical qubits in QLDPC code blocks and connects to processing units through ports supporting Z-type measurements on memory windows.
  • 3. Memory: Memory blocks can be cyclically shifted using local and code-block-scale SWAP operations, avoiding longer-range connectivity requirements.

2. Fully Parallel Operation

The architecture supports fully and partially parallel circuits by joining processing units when entangling operations require coordination and separating them when independent work resumes. Clifford frame cleaning avoids implementing every inter-unit CNOT while allowing compilation to optimize the join–separate schedule.

  • Fully Parallel Operation: Independent circuits can run on separate processing units, reducing logical timesteps to approximately the maximum duration among them.This trades additional qubits for shorter runtime, such as when executing multiple algorithm shots in parallel.
  • Flexibly Parallel Operation: Physical implementation of every inter-unit CNOT can erase parallelism benefits because its time cost scales with the number of entangling gates.This is especially problematic when a highly parallelizable circuit region coexists with a poorly parallelizable one.
  • Flexibly Parallel Operation: Clifford frame cleaning physically applies a Clifford that acts trivially on a chosen subset of logical qubits, using at most 4|K′| logical Pauli product measurements.The construction generalizes an earlier surface-code caching technique to arbitrary generalized-surgery architectures.
  • Flexibly Parallel Operation: Processing units can be separated when parallelism is beneficial, with the choice and timing optimized for each circuit.This can save time relative to fully serial execution or physically implementing all inter-unit entangling gates.
  • Flexibly Parallel Operation: Entangling CNOTs join processing units, after which logical measurements across the joined unit are serialized.A later Clifford frame cleaning step can restore separability using up to 4k additional logical measurements.

4. General Operation

The general architecture combines processing units with optional memory and modular code constructions. Generalised bicycle codes, gadget systems, and bounded-range bridges provide logical measurements, scalable connectivity, and configurable parallel operation.

  • General Operation: Read-only memory access fans data from a memory port into ancillary logical qubits of a processing unit through logical CNOT gates.Memory ports are joined to processing units during access using the same joining and separating concepts as flexible parallelism.
  • General Operation: Arbitrarily large processing units can use physical connections whose scale remains constant in the number of logical qubits.In a two-dimensional arrangement, the connection scale is approximately √npb, the square root of the processing-block size.
  • General Operation: Bounded-range connectivity avoids logical-qubit routing and confines changes between logical cycles to processing-block scale.This supports hardware platforms whose interaction fidelity decreases continuously with distance.
  • General Operation: The architecture’s modular structure supports co-design by associating hardware modules with processing units.Joined modules must be connected in the architecture so logical measurements can span them without long-distance transport.
  • Code Construction: The GB-code instantiation uses weight-six parity checks and short-distance syndrome-extraction transport patterns.The family is parameterized by m > 3 and uses lift l = 2^m − 1, with explicit instances presented for the first five codes.
  • Code Construction: Processing blocks combine GB code blocks with generalized-surgery gadgets that enable selected logical Pauli measurements during error correction.Duplicate gadgets can measure commuting sets of logical operators in parallel.

B. Modules

The module designs instantiate processing and magic-engine components using generalised bicycle codes and ancillary gadgets. Their resource choices balance logical protection, magic-state throughput, and physical-qubit overhead.

  • Processing Blocks: 14β logical qubits require 860β physical qubits at code distance d = 16, while 16β logical qubits require 1620β physical qubits at d = 24.The two configurations provide different protection and capacity trade-offs within the same GB-code family.
  • Magic Engines: Magic engines use 15-to-1 distillation with injected noisy |T⟩ states and post-selection measurements to produce encoded high-fidelity magic states.The protocol uses fifteen Z-type π/8 rotations followed by four logical measurements for post-selection.
  • Magic Engines: At most ten logical measurements are required in parallel because the fifteen rotations are split into two batches and injection measurements are included.The batch structure limits simultaneous state injections to eight.
  • Magic Engines: The magic-engine overhead is nme = ncb + 10ng + 30(na + da −1) + nα.It accounts for the GB code block, ten measurement gadgets, fifteen ancillary codes with bridges, and optional ancillary qubits.
  • Magic Engines: The output-infidelity design chooses GB-code distance de so its logical errors are negligible relative to the target, with rotation errors approximated by prot ≈ pin + (da + 1)pa.The reject probability pα captures failures in preparing sufficiently many noisy magic states for cultivation or post-selection.
  • Magic Engines: For both p = 10^-3 magic-engine cases, the estimated reject rate is pr ≈ 10%.Distillation completes within the d = 24 logical-cycle budget for the p = 10^-3 engine.
  • Magic Engines: Magic-engine times satisfy tme ≤26tc for p = 10^-3 and tme ≤18tc for p = 10^-4.These bounds fit within d = 24 and d = 16 GB logical cycles, except for the stated Fermi-Hubbard p = 10^-4 case.

3. Memory

The optional memory uses the same code blocks as processing units and exposes windows through ports. Its resource formula supports parallel access by multiple processing units while preserving low-overhead storage.

  • Memory: Memory windows match the k/2 logical qubits in a logical sector, with each port implemented by a Z-type gadget and a bridge.Memory-access logical operators commute on every physical qubit and can therefore be measured in parallel.
  • Memory: Memory blocks with at most two ports increase check weight and qubit degree by at most two.The table caption identifies fitted parameters for memory and logical-measurement experiments, including 95% confidence intervals.
  • Memory: 14ν logical memory qubits require 508ν + 88ρ physical qubits at d = 16, while 16ν logical qubits require 1020ν + 150ρ at d = 24.Here ρ processing units access the memory in parallel, and the per-port additions are 88 and 150 physical qubits respectively.

C. Simulation Results

The simulations estimate logical error rates for memory and generalised-surgery measurements across GB code distances and physical error rates. These results support code-distance selection for resource estimation, while fast real-time decoding remains outside scope.

  • Simulation setup: The simulations evaluate memory and generalised-surgery logical measurements over one logical cycle under circuit-level depolarising noise.Memory circuits use d syndrome-extraction rounds with failure rates rescaled by (d + 2)/d; surgery circuits use d + 2 rounds.
  • Simulation results: The ansatz uses distance-independent parameters, enabling extrapolation of collected results across the GB code family.Figure 6 contrasts solid memory-data fits with dashed logical-measurement-data fits and shows 99% confidence intervals for memory points.
  • Interpretation: The simulations benchmark architecture capabilities and guide code-distance choice for resource estimation.Most-likely error decoding correctly decodes all faults of weight less than d/2 and avoids error floors associated with some alternatives.
  • Limitations: Developing a sufficiently fast decoder for real-time use by quantum-hardware classical control is outside the paper’s scope.The authors identify this as future work rather than resolving it through the reported simulations.

2. Implementation and Results

The factoring implementation combines residue-number-system arithmetic with parallel processing of primes, trading a smaller space increase for substantial time savings. The resulting architecture supports low-overhead factoring and related application resource estimates.

  • RSA algorithm: Residue-number-system arithmetic reduces the modular-exponentiation working register from Θ(log N_RSA) to Θ(log log N_RSA) logical qubits.Gidney’s algorithm processes the |P| residue-system primes serially, so this space reduction does not by itself reduce overall runtime.
  • RSA algorithm: The parallelised algorithm achieves orders-of-magnitude time reductions with a smaller increase in space overhead.A single input register can be reused while working registers are duplicated, producing significant spacetime savings.

2. Implementation on Pinnacle Architecture

The RSA implementation maps parallel working registers onto Pinnacle processing units while sharing input memory when possible. Resource accounting combines processing units, memory, ports, magic engines, logical cycles, and expected shots.

  • Processing-unit mapping: Each parallel working register receives a processing unit, which runs independently except during accumulator aggregation.Clifford-frame cleaning after pairwise interactions prevents processing units from remaining joined.
  • Processing-unit mapping: Each working register contains accumulator, discrete-log, ancillary, and Toffoli-compilation sub-registers requiring κ logical qubits.κ = f + 2ℓ + len(m) + 2 max(f, ℓ + len(m)) + 1.
  • Memory and capacity: Allocating ⌈κ/k⌉ processing blocks per unit provides the required logical capacity, while input registers are associated with shared architecture memory.Memory access uses w1-qubit windows for lookup operations targeted at working registers.
  • Resource accounting: The total physical-qubit count is the sum of working-register and memory components, ntotal = nw + nm.Memory and ports contribute additional physical qubits beyond the processing units and magic engine.
  • Runtime accounting: Runtime per shot is t = Tt_l, where the logical-cycle time is t_l = d_t t_c, and total factoring runtime is t_total = σt.The expected shot count depends on the Ekera-Håstad parameter and the probability that a shot has no logical error.

4. Results

Resource optimisation evaluates RSA-2048 factoring across physical error rates, code-cycle times, and parallelisation choices. Under the stated assumptions, Pinnacle reaches sub-million-qubit factoring and reports several runtime–qubit trade-offs.

  • Optimisation: The optimisation varies Gidney-algorithm parameters and the parallelisation factor over 1 ≤ ρ ≤ |P| under prime-availability and window-size constraints.The results are reported in Table VI and Fig. 3.
  • RSA-2048 results: Fewer than one hundred thousand physical qubits factor RSA-2048 at p = 10^-3 in an expected runtime of one month.With the same error rate and code cycle time, 139 thousand qubits achieves one week and 400 thousand achieves one day.
  • RSA-2048 results: With one million physical qubits, factoring takes an expected eight hours, compared with five days in Ref. [8].This reports the runtime comparison at the same physical error rate and code cycle time.
  • Alternative hardware regimes: At p = 10^-4, the minimum physical-qubit requirement is 53 thousand.For a 1 ms trapped-ion code cycle, one million physical qubits yields an expected factoring time below three months.
  • Alternative hardware regimes: At p = 10^-3 and a 1 ms code cycle, 19 million qubits yield a 15-day runtime, versus 5.6 days in Ref. [17].The comparison concerns a neutral-atom platform; Ref. [17] uses transversal gates with algorithmic fault tolerance to reduce logical-cycle time.

VII. CONCLUSION

The Pinnacle Architecture uses QLDPC codes to reduce the physical-qubit overhead of universal quantum computing. It factors 2048-bit RSA integers with fewer than one hundred thousand physical qubits, while further progress remains possible through higher-rate codes and component optimisation.

  • VII. CONCLUSION: Fewer than one hundred thousand physical qubits suffice to factor 2048-bit RSA integers on the Pinnacle Architecture.The conclusion contrasts this with close to a million physical qubits for surface-code architectures.
  • VII. CONCLUSION: QLDPC codes provide order-of-magnitude overhead reductions compared with surface-code architectures.The architecture leverages the high encoding rate of QLDPC codes for universal quantum computing.
  • VII. CONCLUSION: Scaling from one hundred thousand to one million physical qubits poses challenges including networking between separated devices on many hardware platforms.The smaller target device size could therefore hasten the onset of practical quantum computing.
  • VII. CONCLUSION: Higher-rate QLDPC codes and further component optimisation could plausibly achieve the same order-of-magnitude reduction within the Pinnacle Architecture.The generalised bicycle codes used in this work are not the highest-rate QLDPC codes known.

Appendix A: Cost of Clifford Frame Cleaning

The appendix establishes the Pauli-rotation costs required for Clifford frame cleaning. It proves a general 4w-step construction and a 2w-step construction for cleaning a memory port by mapping Clifford operators to support on the remaining qubits.

  • Formalism: A Pauli operator is represented by a vector in Z2n, with commutation determined by the symplectic inner product.Vectors with inner product 0 commute, while those with inner product 1 anti-commute.
  • Formalism: A Pauli π/4 rotation maps an anti-commuting Pauli operator by multiplying it with the rotation axis and leaves commuting operators unchanged.In vector form, its action is Eu(v) = v + ⟨u, v⟩u.
  • General Clifford frame cleaning: 4w Pauli π/4 rotations can clean w qubits from an arbitrary n-qubit Clifford operator, leaving support only on the last n − w qubits.The proof proceeds inductively, using four rotations at each step to map successive rows into the desired form.
  • Memory-port cleaning: 2w Pauli π/4 rotations suffice to clean a memory port when the Clifford operator acts trivially or as the control of a CNOT on the first w qubits.The specialised construction preserves the relevant structure while mapping the operator to support on the last n − w qubits.
Loading 2602.11457v2…