Source-linked AI summary

Benchmarking a trapped-ion quantum computer with 30 qubits

Jwo-Sy Chen, Erik Nielsen, Matthew Ebert, Volkan Inlek, Kenneth Wright, Vandiver Chaplin, Andrii Maksymov, Eduardo Páez, Amrit Poudel, Peter Maunz, John Gamble

arXiv:2308.05071v2quant-ph

TL;DR

As trapped-ion processors grow larger, it becomes harder to characterize both their component quality and holistic user-facing performance. The paper benchmarks IonQ Forte, a 30-qubit all-to-all system, combines component and application measurements, and models application circuits from component data. Forte reaches #AQ 29, while simulations correlate with experiments but generally overpredict application performance because of errors omitted from the simple model.

  • Problem

    Larger many-qubit processors create challenges in calibrating gate pairs and judging holistic performance, requiring both component-level and application-oriented characterization.

  • Method

    The study benchmarks all 435 qubit pairs and application circuits on IonQ Forte, then simulates the applications with a depolarizing model derived from measured component-level error rates.

  • Results

    The system passes the application-oriented benchmarks through #AQ 29, while component-based simulations correlate with experiment but predict higher performance than observed.

  • Takeaways & Limitations

    Combining component-level and application-oriented benchmarks provides a broader characterization of quantum-computer performance than either benchmark type alone.

  • Takeaways & Limitations

    The depolarizing model omits error structure and fails to reproduce measured application performance across most tested circuits, with predictions exceeding observations.

Abstract

from arXiv · show

Quantum computers are rapidly becoming more capable, with dramatic increases in both qubit count and quality. Among different hardware approaches, trapped-ion quantum processors are a leading technology for quantum computing, with established high-fidelity operations and architectures with promising scaling. Here, we demonstrate and thoroughly benchmark the IonQ Forte system: configured as a single-chain 30-qubit trapped-ion quantum computer with all-to-all operations. We assess the performance of our quantum computer operation at the component level via direct randomized benchmarking (DRB) across all 30 choose 2 = 435 gate pairs. We then show the results of application-oriented benchmarks and show that the system passes the suite of algorithmic qubit (AQ) benchmarks up to #AQ 29. Finally, we use our component-level benchmarking to build a system-level model to predict the application benchmarking data through direct simulation. While we find that the system-level model correlates with the experiment in predicting application circuit performance, we note quantitative discrepancies indicating significant out-of-model errors, leading to higher predicted performance than what is observed. This highlights that as quantum computers move toward larger and higher-quality devices, characterization becomes more challenging, suggesting future work required to push performance further.

1 Introduction

The paper addresses how to characterize increasingly complex many-qubit quantum processors by combining component-level and application-oriented benchmarks. It benchmarks a 30-qubit all-to-all trapped-ion system, reaches #AQ 29, and tests how well component measurements predict application performance.

  • 1 Introduction: Component-level benchmarks diagnose hardware behavior, whereas application-oriented benchmarks better capture end-to-end user experience including compilation and error mitigation.The paper argues that both benchmark types are needed because each exposes different aspects of quantum-computer performance.
  • 1 Introduction: The study connects component measurements to application benchmarks by simulating application circuits with a QPU model derived from component-level results.The comparison tests whether component-level performance can predict application-sized circuit behavior.
  • 1 Introduction: The 30-qubit IonQ Forte processor has all-to-all connectivity in a single linear ion chain, enabling benchmarking of all 435 possible qubit pairs.The study uses direct randomized benchmarking to assess component-level performance across every pair.
  • 1 Introduction: #AQ 29 is achieved, with acceptable outcomes for representative circuits containing up to 841 pre-optimized two-qubit gates.The threshold is defined as greater than 1/e circuit Hellinger fidelity.
  • 1 Introduction: The paper is organized around system setup, component and application benchmarking, QPU-model simulation, and comparison between simulated and experimental results.This structure includes compiler optimization and error-mitigation effects in the application-oriented analysis.

2 Experimental setup

IonQ Forte is a cryogenic, surface-trap processor using individually addressable Raman beams, amplitude-modulated entangling gates, and automated control software. Its 30-qubit configuration supports all-to-all operations, while compilation and circuit variants incorporate hardware mapping and error-mitigation strategies.

  • 2 Experimental setup: Forte uses a surface linear Paul trap with separate loading and quantum-operation zones, forming a 36-ion chain configured here as 30 qubits.The cryostat reaches temperatures below 10 K, and ions are transported from loading into the operation zone.
  • 2 Experimental setup: Ion imaging uses an in-vacuum high-NA lens, an individual multimode fiber for each ion, and PMTs for simultaneous photon counting.The system optically pumps ions before gates and reads out the quantum state through fluorescence detection.
  • 2 Experimental setup: Four AODs steer two counter-propagating Raman beam pairs along the trap axis, allowing independent alignment of beams to individual ions.Continuous AOD frequency tuning reduces alignment errors and supports flexible ion spacing.
  • 2 Experimental setup: The arbitrary-angle XXij(χ) gate uses Mølmer–Sørensen amplitude-modulated pulses, with χ specifying entanglement angle and i,j identifying the qubits.Single-qubit π/2 gates surround the entangling operation to form a phase-insensitive ZZij(χ) gate.
  • 2 Experimental setup: A software runtime compiles circuits to native gates, remaps circuit qubits to physical ions, and maintains pulse parameters through calibration updates.Each circuit can be compiled into 25 equivalent variants with different decompositions and physical-qubit assignments for error mitigation by symmetrization.

3 Benchmarking

Forte is benchmarked at both the component and application levels, combining direct randomized benchmarking with application-circuit fidelity and #AQ evaluation. The system shows low single-qubit error, broader two-qubit variability, no significant distance correlation, and application performance reaching #AQ 29.

  • Benchmarking rationale: Application-oriented benchmarks complement component tests because they measure whole algorithms, whereas component metrics provide more specific information about localized hardware errors.The paper therefore uses both benchmark types and compares component-based expectations with application behavior.
  • Component-level benchmarking: Single-qubit DRB error rates are centered at a median of 2.0 × 10−4, with the 10th–90th percentile spanning 1.8 × 10−4 to 2.6 × 10−4.Each qubit was benchmarked repeatedly using randomized π/2 rotations around the x and y axes.
  • Component-level benchmarking: Two-qubit DRB error rates have a median of 46.4 × 10−4, most pairs fall between 35–100 × 10−4, and the worst pair reaches 885 × 10−4.The measurements cover all qubit pairs, with at least four runs per pair; reported error bars are 20–40% of the values.
  • Component-level benchmarking: Two-qubit infidelity shows no statistically significant correlation with ion distance, with fitted slope 0.17(45).For Forte’s long ion chain, the authors conclude that chain length does not currently limit gate infidelity.
  • Application-oriented benchmarking: Application fidelity is computed from measured and ideal outcome distributions after aggregating compiled circuit variants, while #AQ additionally includes compiler and error-mitigation gains.Figure 6 uses simple aggregation, whereas the #AQ evaluation uses error-mitigated distributions and plurality voting.
  • Application-oriented benchmarking: Forte achieves #AQ=29, with circuits up to width 29 and pre-optimized two-qubit gate count 29^2 = 841 meeting the success threshold.The benchmark incorporates compiler optimization and error mitigation, and the #AQ limit is depth- rather than width-limited in the reported results.

4 Simulation of application circuits

The paper tests whether a depolarizing model built from component-level benchmarks can predict application-circuit performance. Simulations correlate with observations but systematically overestimate fidelity, revealing substantial out-of-model errors and algorithm-dependent discrepancies.

  • Model scope: The model intentionally omits structured component-level errors and is used to test whether additional effects such as drift or crosstalk matter.The paper defers more detailed error modeling to future work.
  • Noise model: The model assigns identical depolarization rates to all single- and two-qubit gates, with rates derived from median DRB error distributions.The model uses ϵ1Q = 2.0 × 10−4 and ϵ2Q = 46.4 × 10−4.
  • Simulation method: The simulations compare noisy executions of the compiled circuit variants with observed and ideal outcome distributions using fidelity.Variant circuits are simulated with Qsim and compared through fidelity to the ideal distribution.
  • Simulation results: The component-level depolarizing model correlates with experiment but predicts higher application-circuit fidelity than observed, indicating significant out-of-model errors.Observed circuits suggest an error rate near 100 × 10−4, whereas the model uses ϵ2Q = 46.4 × 10−4.
  • Simulation results: Fidelity differences do not scale solely with circuit size; circuit structure and algorithm class substantially affect model accuracy.The volumetric comparison shows algorithm-dependent deviations, with larger Hamiltonian-simulation circuits inferred to have the largest differences.

5 Conclusion

IonQ Forte demonstrates a 30-qubit single-chain trapped-ion processor with high-fidelity operations and application performance reaching #AQ 29. However, component-level simulations overpredict application performance, motivating system-level characterization of context-dependent noise.

  • Component benchmarks: The 30-qubit single-chain processor supports high-fidelity operations without appreciable performance degradation from ion separation.The system was exhaustively benchmarked across its one- and two-qubit gates.
  • Application benchmarks: Error mitigation suppresses noise enough for every circuit within the #AQ 29 benchmark to exceed the 1/e fidelity threshold.Before mitigation, useful depth typically reaches about 200 two-qubit gates, with depth varying by algorithm.
  • System-level modeling: Component-level depolarization simulations correlate with experiments but fail to reproduce most application-benchmark results, overestimating observed performance.The authors hypothesize stochastic spin-phase noise transferred through optical path-length fluctuations and producing context dependence.
  • Outlook: The conclusion emphasizes system-level characterization tools as quantum processors become harder to understand from constituent components alone.The paper connects this need to increasingly important interactions among many qubits.

7 Data availability

The paper makes its experimental and simulation materials available through a supplemental data repository.

  • Data availability: The supplemental repository contains the paper’s data, circuit variants, raw experimental and simulation results, and code for error mitigation and figure generation.

A Direct randomized benchmarking

Direct randomized benchmarking estimates single- and two-qubit error rates from exponential decay of randomized circuit success probabilities. The protocol uses separate two-qubit gate-sampling ratios and deeper single-qubit or selected-pair experiments to assess fit stability.

  • Protocol: DRB fits success probability versus circuit depth to a + b p^d, using p as the primary estimate of average gate quality.Single-qubit DRB samples four circuits at depths 1, 10, 100, and 1000, with 100 repetitions each.
  • Two-qubit estimation: Two-qubit DRB uses p2Q values of 0.25 and 0.75 to separate single-qubit and two-qubit error rates from the resulting decay rates.The extracted rates are denoted r1Q and r2Q.
  • Precision and depth: The choice of four circuits and maximum depth 100 balances runtime and precision, with bootstrapped error bars reaching 20–40% of reported two-qubit error rates.Selected pairs were also tested to depth 1000.
  • Fit validation: Deep and shallow fits differ by less than 10%, consistent with bootstrap uncertainty, and the exponential fit supports a time-independent depolarizing description for the tested data.A substantial fidelity drift over time would not be consistent with this interpretation.

B Plurality voting procedure

Plurality voting aggregates equivalent noisy circuit variants into a single histogram through a two-step procedure.

  • B Plurality voting procedure: Plurality voting first divides the total shots among variants implemented through different qubit assignments or compilation strategies.Without noise, the variants represent equivalent circuits.
  • B Plurality voting procedure: The resulting variant histograms are then post-processed by a voting procedure to form one aggregate histogram.

1. Choose a threshold t

The procedure selects bit strings appearing across enough variants, aggregates accepted outcomes, and lowers the threshold when consensus is insufficient.

  • 1. Choose a threshold t: For each randomized shot row, the plurality bit string is added to the aggregate histogram only when its occurrences exceed threshold t.The process repeats after re-randomization until the aggregate histogram converges.
  • 1. Choose a threshold t: If no bit strings are accepted, t decreases by one; at t = 2, counts are averaged across shots instead of using plurality voting.The fallback indicates insufficient consensus among variants for plurality voting.
  • 1. Choose a threshold t: The implementation uses 25 variants, 100 shots per variant, and an initial threshold of t = 7, selected heuristically through initial system testing.
  • 1. Choose a threshold t: The output probability of a bit string is its probability of appearing in at least t variants without another bit string appearing in more variants simultaneously.The method avoids inefficient brute-force sampling when high-threshold pluralities are rare.
  • 1. Choose a threshold t: The algorithm retains bit strings appearing in at least t variants and removes variants containing no retained bit strings.This reduces the collection analyzed for subsequent probability calculations.
  • 1. Choose a threshold t: For t ≥ Nv/2, the algorithm computes exact majority-vote probabilities; for t < Nv/2, it neglects rare competing-bit-string events and is approximate.The approximation could be corrected with inclusion-exclusion, but that is omitted for computational simplicity.

C Mølmer–Sørensen gate properties

Forte’s entangling gates use per-pair amplitude-modulated pulses optimized across multiple motional modes, producing gate durations constrained by mode participation and hardware limits.

  • C Mølmer–Sørensen gate properties: Forte’s amplitude-modulated entangling pulses are optimized per ion pair because closely spaced motional modes contribute to gate infidelity.Neighboring modes are separated by approximately 10 kHz, and nonclosed phase-space trajectories create infidelities.
  • C Mølmer–Sørensen gate properties: A numerical search varies pulse modulation and detuning to close phase-space trajectories while satisfying infidelity, motional-frequency-stability, and optical-power constraints.
  • C Mølmer–Sørensen gate properties: The 435 two-qubit ion-pair gates are characterized by distributions of detunings and MS-gate durations.The detuning histogram is compared with motional-mode frequencies, including the center-of-mass mode.
  • C Mølmer–Sørensen gate properties: 550–883 µs MS-gate durations have a 672 µs median, yielding ZZ-gate durations of 770–1103 µs.
Loading 2308.05071v2…