Source-linked AI summary
Architectural considerations in the design of a superconducting quantum annealing processor
P. I. Bunyk, E. Hoskinson, M. W. Johnson, E. Tolkacheva, F. Altomare, A. J. Berkley, R. Harris, J. P. Hilton, T. Lanting, J. Whittaker
TL;DR
The paper addresses how to build a scalable quantum-annealing processor with an algorithmically useful hardware graph and embedded programmable control circuitry. It develops a Chimera-based architecture and ultra-low-power Φ-DAC programming system, successfully operating processors with up to 512 qubits using 56 control lines and about 65 fJ for complete reprogramming.
Problem
The processor must provide non-planar connectivity and support embedding diverse problem graphs while accommodating control circuitry at scale.
Method
The design uses Chimera unit tiles, shortened qubit wiring across six superconducting metal layers, and embedded Φ-DACs with XYZ addressing optimized for flux storage density.
Results
Up to 512 rf-SQUID qubits operated using 56 control lines, while complete reprogramming dissipated about 65 fJ on chip.
Takeaways & Limitations
The architecture combines scalable hardware connectivity with programmable control infrastructure and negligible static power dissipation for practical quantum-annealing processors.
Abstract
from arXiv · showhide
We have developed a quantum annealing processor, based on an array of tunably coupled rf-SQUID flux qubits, fabricated in a superconducting integrated circuit process [1]. Implementing this type of processor at a scale of 512 qubits and 1472 programmable inter-qubit couplers and operating at ~ 20 mK has required attention to a number of considerations that one may ignore at the smaller scale of a few dozen or so devices. Here we discuss some of these considerations, and the delicate balance necessary for the construction of a practical processor that respects the demanding physical requirements imposed by a quantum algorithm. In particular we will review some of the design trade-offs at play in the floor-planning of the physical layout, driven by the desire to have an algorithmically useful set of inter-qubit couplers, and the simultaneous need to embed programmable control circuitry into the processor fabric. In this context we have developed a new ultra-low power embedded superconducting digital-to-analog flux converters (DACs) used to program the processor with zero static power dissipation, optimized to achieve maximum flux storage density per unit area. The 512 single-stage, 3520 two-stage, and 512 three-stage flux-DACs are controlled with an XYZ addressing scheme requiring 56 wires. Our estimate of on-chip dissipated energy for worst-case reprogramming of the whole processor is ~ 65 fJ. Several chips based on this architecture have been fabricated and operated successfully at our facility, as well as two outside facilities (see for example [2]).
I. INTRODUCTION
Scaling superconducting quantum annealing processors requires balancing precise qubit control, useful interactions, readout, thermalization, and embedded control circuitry. D-Wave Two addresses these constraints through zero-static-power XYZ control, denser multilayer fabrication, and reduced qubit dimensions.
- I. INTRODUCTION: Large-scale quantum devices must support precise individual qubit control, computationally interesting interactions, and high-fidelity qubit-state readout.These requirements become important when moving beyond processors containing only a few dozen devices.
- I. INTRODUCTION: D-Wave One’s SFQ demultiplexer used O(log(N)) control lines but generated heat that required about 1 s to dissipate before computation.The processor typically computed in ∼20µs, making the post-programming delay unacceptable.
- I. INTRODUCTION: D-Wave Two eliminated static power dissipation with an XYZ addressing scheme requiring O(3√N) control lines.Although the scaling is not logarithmic, it was considered sufficiently weak for substantially larger processors in the existing apparatus.
- I. INTRODUCTION: Reducing qubit wiring length by a factor of two through six superconducting metal layers increased overall processor density by a factor of four.The design links shorter wiring to increased qubit energy scale under fixed-temperature constraints.
- I. INTRODUCTION: The paper proceeds from whole-chip requirements and hardware-graph selection to bottom-up control-circuit implementation, while deferring readout infrastructure.This organization reflects the need to coordinate topology and embedded control circuitry.
II. CIRCUIT TOPOLOGY
The Chimera hardware graph is designed for optimization problems under physical implementation constraints, combining non-planarity, complete-graph embedding capability, and practical control integration. Its parameters support an Ising formulation with discrete spins and bounded programmable fields and couplings.
- II. CIRCUIT TOPOLOGY: The optimization target is a quadratic form over discrete variables s_i ∈ {−1, +1}, with h_i and J_ij taking values from −1 through +1 in increments of 1/8.The hardware graph G defines the nodes and edges over which the form is minimized.
- II. CIRCUIT TOPOLOGY: Chimera was designed to support optimization problems while satisfying physical implementation constraints.The topology is evaluated against requirements including non-planarity, graph embedding, and on-chip control circuitry.
- II. CIRCUIT TOPOLOGY: Non-planarity supports the intended NP-complete Ising-spin problems and permits chains of qubits to cross each other.The paper identifies both computational and embedding-related motivations for this requirement.
- II. CIRCUIT TOPOLOGY: Embedding complete graphs lets one processor represent diverse problem topologies by mapping logical qubits onto connected physical-qubit subcomponents.The hardware graph is assessed by the variety of problem graphs it can contain as minors and by the largest embeddable complete graph KM.
- II. CIRCUIT TOPOLOGY: The processor minimizes the need for dedicated analog lines by incorporating programmable control circuitry directly into the hardware graph.This requirement becomes increasingly important as the number of controllable qubits and couplers grows.
B. Constraints
The processor layout must trade off qubit connectivity, compact coupled structures, noise and cross-talk suppression, and regularity. Chimera’s repeatable unit-tile organization provides a practical way to scale this constrained geometry across the chip plane.
- B. Constraints: Each qubit can connect to only ≲10 others before non-ideal response and reduced coupling energy scales become problematic.This limited fan-out prevents direct implementation of an arbitrarily large complete graph.
- B. Constraints: Qubit and coupler lengths should be minimized so that their physical extent remains magnetically coupled to connected partners.Long uncoupled sections would undermine qubit energy scales and coupling strengths.
- B. Constraints: Flux qubits and couplers require minimized pickup areas and cross-talk because their rf-SQUID structure is sensitive to magnetic-field disturbances.The cited disturbances include external flux noise and unintended coupling from nearby circuitry.
- B. Constraints: A unit tile is a smaller qubit structure replicated across both dimensions of the chip plane to simplify design and operation.Regularity is preferred over highly irregular arrangements for general-purpose processors.
C. Chimera topology
The Chimera topology balances useful graph embeddings with physical layout and control-circuit constraints. Its tiled, non-planar structure supports complete-graph embeddings while preserving strong coupling and reducing noise and cross-talk.
- Each Chimera unit tile contains eight qubits forming a complete bipartite graph K4,4, with neighboring tiles extending horizontal and vertical couplings.
- The non-planar Chimera graph supports embedding complete graphs up to 4N nodes in an N × N grid of unit cells.
- The topology interleaves qubits, couplers, and Φ-DAC control circuitry, with three Φ-DACs placed in each intersecting plaquette.
- Strong ferromagnetic couplings contract physical qubits into chains, while tunable couplers connect every chain pair in the embedded complete graph.
- 80 physical qubits can form 16 chains in the N = 4 example, embedding K4N=16.
- Long, narrow differential microstrip loops maximize qubit–coupler coupling while minimizing noise and parasitic cross-talk pickup.
- Chimera tiles scale into arbitrarily large 2D structures, with M = 4 chosen because required Φ-DACs fit efficiently in a 5 × 5 plaquette array.
III. DESIGN AND OPERATION OF A Φ-DAC.
The Φ-DAC design uses inductive storage and staged flux division to provide programmable range and precision while limiting area and wiring. Most devices use two stages, with margins for fabrication variation.
- Individual Φ-DACs generally require about 8 bits of dynamic range, with full ranges from several thousandths of mΦ0 to half a Φ0.
- Most Φ-DACs are two-stage devices chosen to achieve required dynamic range while minimizing control-circuit area and programming wires.
- Individual DAC digits store positive or negative flux quanta in SQUID loops, with SFQ pulses adding or subtracting quanta from each storage loop.
- An MSD loop storing up to 8 flux quanta provides 16 output values, implementing a 4-bit DAC when junction-inductance corrections are negligible.
- A second LSD stage with division ratio 16 subdivides each MSD step into 16 values, producing an 8-bit DAC.
- Storage capacity includes margin so the DACs retain total range and LSD coverage despite fabrication variations.
- Two-stage designs are sufficient for almost all DACs, while devices with different digit counts and weights can use the same design principles.
A. Φ-DAC: Inductive storage and ladder
The physical Φ-DAC implementation combines dense superconducting inductors, magnetic coupling, shielding, and three-port modeling. Layout choices accommodate both low-range and high-range target control while accounting for nonideal coupling paths.
- Large storage inductors of approximately 1 nH are implemented as stacked spirals across four metal layers with 0.25 µm line width and spacing.
- The inductive ladder uses two galvanically connected bottom-metal washers magnetically coupled to the coils, with a shared inductance LDIV between them.
- A top-metal shielding sky-plane covers the Φ-DAC structure to reduce unintended coupling between DAC coils and other circuit elements.
- Simple transformer coupling provides several tens of mΦ0 of output flux, whereas high-range DACs merge the target CJJ loop with the MSD stage to reach approximately half a Φ0.
- A special qubit CCJJ major-loop DAC uses 5 bits, while a coarse qubit flux-bias stage handles larger local flux offsets.
- The realistic layout includes direct LSD–MSD and MSD–output coupling paths in addition to the intended inductive-ladder path.
- The complete Φ-DAC is modeled as a three-port device comprising LSD, MSD, and OUT, with its inductance matrix extracted using FastHenry.
B. Φ-DAC: SFQ pulse sources
Φ-DACs use dc-SQUID SFQ pulse sources to add or subtract single flux quanta to storage loops, with bias levels carefully margined to permit intended transitions while suppressing unwanted programming.
- SFQ pulse-source operation: A current-biased dc-SQUID with two shunted junctions serves as the SFQ pulse source for each Φ-DAC.The source feeds the storage loops used by the DAC stages.
- SFQ pulse-source operation: During programming, PWR biases the junctions, ADDR supplies an initial flux bias, and a TRIG ramp drives sequential junction flips that admit one flux quantum.The process returns the source to its zero-flux state while increasing the storage-loop phase by 2π.
- SFQ pulse-source operation: Reversing PWR adds single flux quanta of the opposite magnetic-field direction to the storage loop.Thus the same pulse mechanism supports both addition and subtraction of stored flux.
- Stage selection and biasing: ADDR and TRIG polarity selects whether the LSD or MSD dc-SQUID operates, while PWR determines the pulse sign.The twisted TRIG connection keeps the unselected source quiescent.
- Stage selection and biasing: Margining chooses PWR, ADDR, and TRIG levels so intended three-line transitions occupy active regions and subset-addressed devices remain in forbidden-zone-free conditions.The critical boundary is parameterized by Φb and Ib, and the levels are chosen to maximize active-region size while avoiding unwanted transitions.
C. Φ-DAC reset
Reliable Φ-DAC operation requires resetting every device from an unknown state to zero before programming. Reset de-programs stored flux one quantum at a time, but temporarily violates normal margining because all DACs reset simultaneously.
- Reset protocol: Φ-DACs must be reset to a known state before realistic programming from an unknown initial state.The protocol starts by setting IPWR to zero.
- Reset protocol: Large ADDR+TRIG pulses de-program a nonzero Φ-DAC one SFQ at a time until it reaches the zero-SFQ state.At zero SFQ, the circulating current in the main loop is zero.
- Reset requirements: Reliable arrival at zero SFQ requires the pulse-source junction critical-current difference to remain well under Φ0/L.Here L is the main Φ-DAC loop inductance.
- Reset requirements: Reset violates the margining criteria because all Φ-DACs are reset simultaneously.
D. Minimizing Φ-DAC footprint
The Φ-DAC footprint is minimized by maximizing stored-flux capacity per area while preserving the required L×Ic product. This footprint sets unit-tile size, qubit length, and ultimately qubit energy scale.
- Design objective: Φ-DAC area ultimately determines processor unit-tile size, qubit wiring length, and qubit energy scales.Minimizing control-circuit area therefore supports shorter qubits and higher energy scales.
- Design objective: Maximum stored flux, which determines DAC range and precision, is proportional to the storage-loop L×Ic product.The design problem is to maximize this product within a fixed area.
- Area optimization: For fixed fabrication layers, spiral-coil inductance and junction area scale with physical area, motivating a balanced choice of source-junction Ic and storage-loop L.The source critical current was chosen against storage inductance under fixed junction-current-density constraints.
- Area optimization: A factor-of-6 increase in Ic combined with a factor-of-6 decrease in L reduces total Φ-DAC area by the same factor while preserving L×Ic.Simply replacing junctions with smaller equal-critical-current junctions would save less than a factor of 2 in area, even with a factor-of-36 Jc increase.
E. XYZ-addressing line count
XYZ addressing reduces the wiring needed to program 4608 Φ-DACs by reusing address, trigger, and power lines across a tiled processor layout. The implemented arrangement uses 56 lines and favors regularity over strict optimality.
- Addressing architecture: Each unit tile requires 72 Φ-DACs to control its qubits and couplers, yielding 4608 Φ-DACs across the 512-qubit processor.The tile allocation is 6 DACs per qubit, 16 for internal couplers, and 8 for external couplers.
- Addressing architecture: Within a unit tile, 15 ADDR and 5 TRIG lines select the 72 Φ-DACs arranged in 25 three-DAC plaquettes.One plaquette is empty, and one of three DACs is selected by an ADDR line while sharing a TRIG line.
- Addressing architecture: The 8×8 tile array is divided into sixteen 2×2 PWR domains, with each domain supplying its series-connected Φ-DACs through one PWR line.
- Physical implementation: The active processor portion occupies approximately 3.5×3.5 mm2 and contains an 8×8 array of 8-qubit unit tiles.Each unit tile is 335 µm on a side.
- Addressing architecture: 56 lines address all Φ-DACs: 30 ADDR, 10 TRIG, and 16 PWR lines, with ADDR and TRIG reused between power domains.The arrangement is close to optimal but was selected partly to maintain a more regular layout.
IV. CONCLUSIONS
The processor architecture integrates a 512-qubit hardware graph with embedded control infrastructure, enabling operation using 56 programming lines. Its Φ-DAC design eliminates static power dissipation and reduces post-programming thermalization to 10 ms, a 100-fold improvement over D-Wave One.
- 512 rf-SQUID qubits were successfully operated using only 56 control lines for problem programming.The architecture combines the processor hardware graph with its required control infrastructure.
- Zero static power dissipation is achieved by serially biasing Φ-DAC devices with a fixed current set by a room-temperature resistor.Only energy associated with moving flux quanta into or out of storage is dissipated on chip.
- 65 fJ is the estimated on-chip energy dissipated when all 9216 Φ-DAC stages are completely reprogrammed from -16 to +16 SFQ.A pair of 55 µA Φ-DAC junctions dissipates 0.22 aJ per flux quantum moved.
- 10 ms is the D-Wave Two post-programming thermalization time, a factor of 100 improvement over D-Wave One’s approximately 1 s delay.The processor can return to approximately 20 mK within this interval.