Source-linked AI summary
Strong quantum computational advantage using a superconducting quantum processor
Yulin Wu, Wan-Su Bao, Sirui Cao, Fusheng Chen, Ming-Cheng Chen, Xiawei Chen, Tung-Hsun Chung, Hui Deng, Yajie Du, Daojin Fan, Ming Gong, Cheng Guo, Chu Guo, Shaojun Guo, Lianchen Han, Linyin Hong, He-Liang Huang, Yong-Heng Huo, Liping Li, Na Li, Shaowei Li, Yuan Li, Futian Liang, Chun Lin, Jin Lin, Haoran Qian, Dan Qiao, Hao Rong, Hong Su, Lihua Sun, Liangyuan Wang, Shiyu Wang, Dachao Wu, Yu Xu, Kai Yan, Weifeng Yang, Yang Yang, Yangsen Ye, Jianghan Yin, Chong Ying, Jiale Yu, Chen Zha, Cha Zhang, Haibin Zhang, Kaili Zhang, Yiming Zhang, Han Zhao, Youwei Zhao, Liang Zhou, Qingling Zhu, Chao-Yang Lu, Cheng-Zhi Peng, Xiaobo Zhu, Jian-Wei Pan
TL;DR
Quantum computational advantage requires scalable quantum processors whose performance can outpace improving classical hardware and algorithms. This paper develops the 66-qubit Zuchongzhi processor and benchmarks it with random circuit sampling, reaching 56 qubits and 20 cycles. The resulting task is estimated to require 2–3 orders of magnitude more classical computational cost than the previous 53-qubit Sycamore task and about 8.2 years on Summit versus 1.2 hours experimentally.
Problem
Demonstrating quantum computational advantage requires scaling qubit number and control fidelity so quantum performance can outpace continuing classical hardware and algorithmic improvements.
Method
The paper develops a programmable 66-qubit superconducting processor and evaluates it using random quantum circuit sampling, with classical cost estimated using tensor-network and Schrödinger–Feynman algorithms.
Results
2–3 orders of magnitude higher classical computational cost is estimated for the 56-qubit, 20-cycle task than for the previous 53-qubit, 20-cycle task.
Takeaways & Limitations
56-qubit, 20-cycle sampling is reported as an intractable classical-simulation task and establishes a new challenge to classical computing capability.
Abstract
from arXiv · showhide
Scaling up to a large number of qubits with high-precision control is essential in the demonstrations of quantum computational advantage to exponentially outpace the classical hardware and algorithmic improvements. Here, we develop a two-dimensional programmable superconducting quantum processor, \textit{Zuchongzhi}, which is composed of 66 functional qubits in a tunable coupling architecture. To characterize the performance of the whole system, we perform random quantum circuits sampling for benchmarking, up to a system size of 56 qubits and 20 cycles. The computational cost of the classical simulation of this task is estimated to be 2-3 orders of magnitude higher than the previous work on 53-qubit Sycamore processor [Nature \textbf{574}, 505 (2019)]. We estimate that the sampling task finished by \textit{Zuchongzhi} in about 1.2 hours will take the most powerful supercomputer at least 8 years. Our work establishes an unambiguous quantum computational advantage that is infeasible for classical computation in a reasonable amount of time. The high-precision and programmable quantum computing platform opens a new door to explore novel many-body phenomena and implement complex quantum algorithms.
I. INTRODUCTION
Quantum computational advantage requires quantum devices to complete well-defined tasks overwhelmingly faster than classical computers, while continued classical improvements make hardware scaling necessary. Zuchongzhi addresses this challenge with a 66-qubit processor and demonstrates large-scale random circuit sampling.
- Quantum computational advantage is defined as completing a well-defined task overwhelmingly faster than any classical computer within a reasonable time.
- Recent experiments with 53 superconducting qubits and 76 photons provided strong evidence for quantum computational advantage.
- Increasing qubit number is expected to exponentially outpace classical performance as classical algorithms and hardware continue improving.
- Simultaneously increasing qubit number and high-fidelity gate performance is crucial for NISQ development and surface-code logic-qubit demonstrations.
- Zuchongzhi contains 66 qubits, achieves average single-qubit, two-qubit, and readout performances of 99.86%, 99.41%, and 95.48%, and samples circuits up to 56 qubits and 20 cycles.
- 1.2 hours completes the sampling task, while classical simulation is estimated to require at least an unreasonable amount of time.
II. HIGH-PERFORMANCE QUANTUM PROCESSOR
Zuchongzhi is a two-dimensional, tunable-coupling superconducting processor built from 66 qubits and 110 couplers. Calibration and control procedures produce high-fidelity simultaneous single- and two-qubit operations and readout for benchmarking.
- 66 qubits form an 11-by-6 two-dimensional lattice, with neighboring interactions mediated by tunable couplers and separate readout and control hardware.
- Two sapphire chips are stacked using indium-bump flip-chip bonding, separating the qubit-coupler layer from readout, control, and wiring components.
- All 66 qubits and 110 couplers function properly, while 56 selected qubits are optimized for random circuit sampling and classical-simulation complexity.
- 0.14% average single-qubit Pauli error is achieved during simultaneous operation through frequency optimization and parallel cross-entropy benchmarking.
- The iSWAP-like two-qubit gate uses resonant neighboring qubits and calibrated flux-bias pulses, with parameters optimized through parallel XEB.
- 0.59% average two-qubit Pauli error is achieved when all gates operate simultaneously, after optimizing frequencies and calibrating pulse effects.
- 4.52% average single-qubit readout error is obtained after using separate frequency settings and calibration procedures to reduce readout crosstalk.
III. RANDOM QUANTUM CIRCUIT BENCHMARKING
Random quantum circuits benchmark Zuchongzhi by combining random single-qubit layers with patterned two-qubit layers and using classically tractable variants to estimate full-circuit fidelity. The processor maintains predicted fidelity behavior and demonstrates statistically significant sampling for 56 qubits and 20 cycles.
- Random circuit sampling benchmarks overall processor performance using circuits composed of repeated single-qubit and two-qubit gate layers.
- Each circuit cycle applies random single-qubit gates followed by two-qubit gates arranged in the repeating ABCDCDAB pattern.
- Patch circuits remove a two-qubit-gate slice, while elided circuits remove only part of the gates between patches to make XEB fidelity classically estimable.
- 56-qubit circuits are tested from 12 to 20 cycles using full, patch, and elided circuits, with patch and elided results used to assess full-circuit performance.
- (6.62 ± 0.72) × 10^-4 combined linear XEB fidelity is obtained for ten 56-qubit, 20-cycle circuit instances, rejecting uniform sampling at 9σ.
- Observed fidelities and their decay with qubit number and cycle count match predictions from multiplying individual-operation errors, supporting low error correlation.
IV. COMPUTATIONAL COST ESTIMATION
The paper estimates the classical cost of reproducing its hardest 56-qubit, 20-cycle random-circuit sampling task using tensor-network and Schrödinger–Feynman algorithms. Both estimates place the task 2–3 orders of magnitude beyond the previous 53-qubit benchmark, while more efficient classical simulation remains an expected possibility.
- State-of-the-art tensor-network and Schrödinger–Feynman algorithms are used to estimate the classical cost of simulating the hardest circuits.The target is a 56-qubit random circuit with 20 cycles.
- 1.65×10^20 floating-point operations are estimated for one perfect sample from the 56-qubit, 20-cycle circuit, versus 1.63×10^18 for the previous 53-qubit circuit.
- 5.76 × 10^17 core-hours are estimated for Schrödinger–Feynman simulation of the 56-qubit task, versus 8.90 × 10^13 core-hours for the previous 53-qubit task.
- The new 56-qubit, 20-cycle sampling task has an estimated classical cost 2–3 orders of magnitude greater than the previous 53-qubit, 20-cycle task.The authors describe this as enlarging the gap between quantum-device performance and classical simulation.
- More efficient classical simulation approaches are anticipated, sustaining competition between quantum and classical computing and complicating large-scale benchmarking.
V. CONCLUSION
The paper reports a fully programmable 66-qubit superconducting processor and benchmarks a 56-qubit, 20-cycle random-circuit experiment. The authors state that the processor’s scalable architecture is compatible with surface-code error correction and may support future NISQ applications.
- The work reports the design, fabrication, measurement, and benchmarking of a fully programmable 66-qubit superconducting quantum processor.
- A 56-qubit, 20-cycle random-circuit experiment establishes a new record challenging classical computing capability.
- The processor’s scalable architecture is compatible with surface-code error correction and can serve as a test-bed for fault-tolerant quantum computing.
- The authors expect the large-scale, high-performance processor to support NISQ applications beyond classical computers in the near future.
Supplemental Material for “Strong quantum computational advantage using a superconducting quantum processor”
The supplemental material describes the processor’s two-chip fabrication, cryogenic packaging, and room-temperature control and readout infrastructure. It also details the wiring, amplification, digitization, and electronic channel resources used for operation.
- The processor is built from top and bottom chips fabricated on sapphire, with control and readout circuits distributed across the two-chip platform.
- The packaged processor is mounted in a dilution refrigerator and connected to room-temperature electronics through attenuators, filters, and amplifiers.
- DACs generate control and readout signals at room temperature before attenuation and low-pass filtering in the cryogenic wiring.
- JPA, HEMT, and room-temperature amplifiers increase readout signal strength before digitization and demodulation by ADC modules.
- 330 DAC channels, 11 ADC modules, 11 DC channels, and 34 microwave-source channels support the room-temperature electronics.
1. Basic Calibration
The calibration workflow characterizes coherence, control-line distortion, crosstalk, timing, readout, and gate parameters across the processor. It then optimizes operating frequencies and benchmarks high-fidelity single- and two-qubit gates and readout.
- 1. Basic Calibration: Calibration covers all 66 qubits, 110 couplers, 66 readout resonators, and 11 JPAs before performance optimization proceeds.
- 1. Basic Calibration: The procedure measures T1 and T2, bias-line step responses, XY crosstalk, and synchronization among qubit-drive and bias controls.
- 1. Basic Calibration: Operating frequencies are optimized using coherence and crosstalk data, after which single-qubit, readout, and two-qubit parameters are calibrated in parallel.
- Single-Qubit Gate Calibration: 0.14% average simultaneous XEB Pauli error is obtained for single-qubit gates after pulse and frequency optimization.
- Readout Calibration: 95.23% average readout fidelity is achieved, with simultaneous |0⟩ and |1⟩ identification errors of 3.46% and 6.08%, respectively.
- Readout Calibration: Random-bit-string measurements find a 0.14% average increase in single-qubit identification error and a 0.93 factor reduction in 56-qubit state readout fidelity.The factor is used to correct estimated fidelities in random-circuit benchmarking.
- Two-Qubit Gate Calibration: 0.76% average simultaneous XEB Pauli error is obtained across all 110 couplers for the calibrated iSWAP-like two-qubit gates.The gates use a 32 ns duration including 3 ns pulse rise/fall and 26 ns interaction time.
B. Part2: Fine-tune on 56 qubits
The 56-qubit subset was recalibrated and optimized to balance fidelity, sampling uncertainty, and runtime, enabling 20-cycle random circuit sampling within an acceptable sampling time.
- 56-qubit, 20-cycle random circuit sampling was completed as an intractable task for classical simulation within an acceptable sampling time.
- 1.5% two-qubit XEB Pauli error was used as the threshold for additional SPB-fidelity sweeps across interaction frequencies.Optimization used 70 random circuit instances at fixed depth 20 cycles.
- 4.52% average readout error remained after calibration and optimization.
- 0.14% and 0.59% were the average XEB Pauli errors for simultaneous single-qubit and two-qubit gates, respectively.
IV. SOFTWARE SYSTEM
QOS is a scalable quantum operating system that abstracts hardware, manages resources, and executes operations across a complex superconducting processor. Its actor-based, concurrent, pipelined design supports high-performance calibration and experimentation.
- More than 400 control channels must be precisely managed because the 66-qubit processor is complex, constrained, and susceptible to parameter drift.
- QOS abstracts hardware details, manages resources, and implements quantum operations, with the last function being unique to a quantum operating system.
- The kernel uses the Actor Model to parse quantum operations in parallel across agents representing hardware components.
- QOS separates quantum and classical hardware components into concurrently communicating agents, while QThread supports concurrent experiments.
- A pipelined workflow minimizes software overhead, allowing heavy experiments such as parallel randomized benchmarking on large processors with negligible software cost.
- The Schrödinger–Feynman simulation cost is proportional to (2^n1 + 2^n2)rg, where r is Schmidt rank and g is the number of cross-partition gates.
C. Performance of patch circuits and elided circuits
Patch and elided circuits reduce entanglement so their XEB fidelities can be compared with full circuits as performance estimates. Across 15–56 qubits and 10 cycles, both approximations closely tracked full-circuit fidelity.
- 56-qubit coupler activation patterns specify which qubits may interact simultaneously in each cycle.
- Patch circuits remove all cross-partition two-qubit gates, whereas elided circuits remove only some cut gates during early cycles.
- 1.05 and 1.10 were the average patch-to-full and elided-to-full fidelity ratios, respectively, across verification circuits.
- 8% and 9% were the standard deviations of the patch and elided fidelity ratios, dominated by system fluctuations.
- 15–56 qubits and 10 cycles were covered when comparing patch, elided, and full circuit XEB fidelities.
D. Distribution of bitstring probabilities
The section tests whether bitstring-probability distributions from 56-qubit, 20-cycle circuits agree with theoretical XEB predictions. Combined linear and logarithmic XEB measurements reject uniform sampling, while uncertainty estimates agree with theory.
- Distribution of bitstring probabilities: 56-qubit, 20-cycle circuits were sampled with approximately 1.9 × 10^7 bitstrings from each of 10 instances.Ideal probabilities were calculated for the sampled bitstrings and compared with theoretical linear- and logarithmic-XEB curves.
- Distribution of bitstring probabilities: 6.60 × 10^-4 and 5.80 × 10^-4 were the combined linear and logarithmic XEB fidelities, respectively.The combined data came from ten circuit instances.
- Distribution of bitstring probabilities: 9.7 × 10^-11 was the p-value for the null hypothesis F = 0, which was rejected confidently.The combined Kolmogorov–Smirnov test assessed the sampled distributions against the null hypothesis.
- Statistical uncertainties: (6.62 ± 0.72) × 10^-4 and (5.82 ± 0.92) × 10^-4 were obtained for linear and logarithmic XEB across nine circuits.Inverse-variance weighting was used to estimate the fidelities and statistical uncertainties.
- Statistical uncertainties: 7.2 × 10^-5 and 9.2 × 10^-5 were the predicted statistical uncertainties for linear and logarithmic XEB, respectively, agreeing with experiment.Bootstrap samples were also used to verify the uncertainty estimates.
- Classical simulation: The Schrödinger simulator computes all 2^n amplitudes, and its performance was benchmarked on a 1536 GB, four-CPU server.The benchmark used gate fusion and single-precision arithmetic.
- Classical simulation: 8.24 years was estimated to reproduce the 53-qubit and 56-qubit, 20-cycle results using Summit.The estimate used a tensor-network contraction cost of 6.66 × 10^18 and a measured 833.75-second time per perfect sample.
C. Computational cost estimation of SFA for the sampling task
This section estimates the classical cost of simulating the 56-qubit, 20-cycle sampling task with Schrödinger–Feynman methods. It analyzes path counts, imbalance cuts, and gate-induced speedups to quantify the remaining computational burden.
- SFA path count: 42 gates cross the cut, and the 56-qubit, 20-cycle circuit requires 438 × 2^4 paths at 100% fidelity.The first and last three iSWAP-like gates can be simplified from Schmidt rank 4 to rank 2.
- SFA path count: 1.06 × 10^18 core hours were estimated to simulate the circuit at 0.0662% fidelity.The estimate was based on measured execution times for prefix values on a 1536 GB, four-CPU server.
- Imbalance cut: 5.76 × 10^17 core hours was the imbalance-cut estimate, trading lower runtime for greater storage use.The allowed partition imbalance was n1 − n2 ≤ 20, producing partitions of 32 and 24 qubits.
- Comparison with Sycamore: 6466 times was the estimated SFA computational-cost ratio between the 56-qubit and 53-qubit, 20-cycle circuits.The comparison includes the estimated speedup from DCD formation.
- Imbalanced gates: SFA reduces the required paths by selecting the top S paths at target fidelity F rather than calculating all 4^g paths.For balanced gates, the number of paths is F × 4^g; imbalanced gates provide the corresponding speedup.
- Imbalanced gates: |δθ| ≈ 0.036 radians produced a speedup well below one order of magnitude for n = 56, m = 20, F = 0.0662%, and g = 42.The estimate assumes the experimentally observed iSWAP-like gate imbalance.
E. Quantum runtime advantage region
This section estimates quantum runtime advantage by comparing quantum sampling time with classical SA and SFA simulation costs under computational and memory constraints. The estimated advantage region expands rapidly as quantum error rates decline.
- Tensor-network algorithms may be more efficient at low depths, whereas SFA is described as the most efficient current method for large circuits with high depth.The estimate is presented as an illustration of the importance of improving quantum-gate and readout fidelity, while other classical simulators may perform better.
- 51 qubits is the estimated maximum for state-vector simulation when supercomputer memory is below 3 PB.SA requires 2^n+1 bytes to store the complex state vector.
- SFA runtime depends on the number of paths, patch-simulation time, circuit fidelity, and the number of circuit partitions.The runtime is optimized subject to memory constraints and patch-count choices.
- At least 1/F^2 quantum samples are required to keep the standard deviation no greater than the fidelity.The quantum runtime is proportional to the number of samples, with sampling rate C_QC = 1/230 MHz.
- The quantum advantage region enlarges rapidly when error rates decline.The comparison uses fitted classical-runtime constants, optimized SFA runtime, and a memory constraint; Fig. S18 maps the region over circuit depth and qubit number.