Source-linked AI summary
Beyond Moore's technologies: operation principles of a superconductor alternative
I. I. Soloviev, N. V. Klenov, S. V. Bakurskiy, M. Yu. Kupriyanov, A. L. Gudkov, A. S. Sidorenko
TL;DR
High-performance computing faces severe energy-efficiency challenges as semiconductor scaling and thermal constraints weaken. This paper reviews superconducting logic and cryogenic memory, analyzes their computer-circuit design issues, and outlines research directions. It presents superconducting technology as a promising post-Moore alternative, with reported energy-efficient operation and demonstrated circuit approaches, while restricting the review to common solutions rather than comprehensive coverage.
Problem
High-performance computing needs more energy-efficient technologies because semiconductor scaling and thermal constraints increasingly limit performance and power efficiency.
Method
The paper reviews the operating principles and evolution of superconducting logic and cryogenic memory, then analyzes their computer-circuit design issues and research directions.
Results
A tested on-chip DC-to-AC converter provided 4.4 GHz oscillation, and adjustable AC bias enabled circuits based on different logics on one chip.
Takeaways & Limitations
Superconductor digital technology is presented as a promising post-Moore candidate for energy-efficient supercomputing, with logic operation reported at approximately 5–50 GHz and 10^-19–10^-20 J per bit.
Takeaways & Limitations
The review is not comprehensive and considers only the most common superconducting logic and memory solutions.
Abstract
from arXiv · showhide
The predictions of Moore's law are considered by experts to be valid until 2020 giving rise to "post-Moore's" technologies afterwards. Energy efficiency is one of the major challenges in high-performance computing that should be answered. Superconductor digital technology is a promising post-Moore's alternative for the development of supercomputers. In this paper, we consider operation principles of an energy-efficient superconductor logic and memory circuits with a short retrospective review of their evolution. We analyze their shortcomings in respect to computer circuits design. Possible ways of further research are outlined.
Introduction
The paper reviews superconducting logic and memory as post-Moore alternatives motivated by worsening energy constraints in high-performance computing. It surveys their operating principles, design issues, evolution, challenges, and research directions.
- Motivation: Energy efficiency has become a crucial constraint on progress in high-performance computing as semiconductor scaling and thermal limits intensify.The paper links these constraints to reduced integration gains, limited clock frequency, and the Dark Silicon problem.
- Motivation: 15.4 MW is the reported power consumption of Sunway TaihuLight at 93 petaflops, while exaflops systems are predicted to require sub-GW power.Such demand is described as comparable to a small powerplant and associated with very high operating costs.
- Motivation: 20 MW is the roadmap target for exaflops supercomputer power, corresponding to 20 pJ/flop or 50 Gflops/W.Sunway TaihuLight achieves 6 Gflops/W, approximately an order below the stated requirement.
- Motivation: Superconductor digital technology is presented as a promising post-Moore candidate because its switching energy is about 10^-19 J and signal transfer has no power penalty.A prospective study reports up to two orders of magnitude higher energy efficiency than semiconductor counterparts, reaching 250 Gflops/W.
- Scope and approach: The review examines superconducting logic branches, including SFQ and adiabatic logic, alongside four cryogenic memory approaches.It also analyzes computer-circuit design issues and discusses challenges and possible directions for further research.
A. Physical basis underlying logic circuits
Superconducting logic combines superconductivity, magnetic-flux quantization, and Josephson effects to represent and transmit information through quantized flux and voltage pulses. Josephson-junction dynamics provide the nonlinear switching basis, while circuit speed and density remain constrained by device and wiring parameters.
- Physical principles: Superconductivity enables ballistic signal transfer without charging interconnect capacitance, supporting picosecond waveforms over long distances with low crosstalk.Superconducting microstrip lines can approach the speed of light over distances exceeding typical chip dimensions.
- Physical principles: Magnetic flux in a superconducting loop is quantized as Φ = nΦ0, allowing SFQ presence or absence to encode logical 1 or 0.This representation physically localizes information in superconducting circuits.
- Josephson switching: A Josephson junction switches from superconducting to resistive state when current exceeds its critical current, changing loop flux and enabling digital logic.The junction is the nonlinear element and is commonly implemented as an SIS weak link.
- Junction dynamics: The resistively shunted junction model maps junction dynamics to a damped pendulum, with capacitance, resistance, and bias current corresponding to inertia, damping, and applied torque.The Stewart-McCumber parameter βc captures capacitance impact, and βc ≈ 1 helps plasma oscillations decay before the next switch.
- Josephson switching: A 2π increase in Josephson phase produces a voltage pulse with V dt = Φ0, so one junction switch transmits an SFQ pulse.The typical switching energy is approximately 2 × 10^-19 J for Ic ≈ 0.1 mA at 4.2 K.
- Implementation constraints: Practical circuits remain constrained by junction area, wiring dimensions, junction density, and clock frequency.The cited Nb-based technology uses wiring features of about 0.5–1 µm, while circuit complexity is limited to about 2.5 million junctions/cm2.
1. SFQ circuit basic principles of operation
SFQ circuits process information by storing and propagating quantized magnetic-flux pulses through Josephson-junction networks. RSFQ provides a sequential logic basis, but its conventional resistive bias supply creates substantial stationary power dissipation.
- An RSFQ data bus is a Josephson transmission line where successive junction switchings transfer SFQ pulses through superconducting loops.Switching sums the SFQ circulating current with applied bias current, redistributing the pulse toward the next junction.
- In RSFQ, an arriving SFQ pulse represents binary “1”, while its absence represents “0”; clocked cells use a decision making pair for readout.The logic cell can operate as a D flip-flop and therefore forms sequential rather than combinational logic.
- RSFQ circuits use clocked operations and were widely adopted for digital and mixed-signal devices after the 1990s.The cited examples include analog-to-digital converters, digital signal processors, and data processors.
- The conventional RSFQ supply uses resistive biasing, with stationary dissipation PS = IbVb and a representative value of ∼800 nW.The feed-line resistors remained a source of stationary energy dissipation even after other resistive connections were replaced.
- At a typical clock frequency of 20 GHz, dynamic dissipation is ∼13 nW, about 60 times lower than stationary dissipation.Reducing stationary dissipation therefore became the main target of RSFQ energy-efficiency improvements.
3. LV-RSFQ
LV-RSFQ reduces bias voltage variation with added inductances, but the approach trades energy savings against clock-frequency limits and circuit area. ERSFQ and eSFQ remove resistive biasing through different phase-balancing strategies, while increasing junction count or imposing design changes.
- LV-RSFQ: LV-RSFQ introduces series inductances with feed-line bias resistors to damp bias-current redistribution and reduce bias voltage.The added inductances limit current variation between neighboring cells.
- LV-RSFQ: Higher clock frequency in LV-RSFQ increases cell voltage, decreases bias current, and can cause cell malfunction.The approach also requires additional feed-line inductance area and does not eliminate static power dissipation.
- ERSFQ: ERSFQ replaces feed-line resistors with Josephson junctions and limiting inductances, enabling circuits to remain in a pure superconducting state.The inductance limits bias-current changes caused by asynchronous junction switching.
- ERSFQ: In ERSFQ, asynchronous switching can produce circulating currents between neighboring cells because their average voltages and Josephson phase increments differ.These currents can be added to the bias current and prevent correct circuit operation.
- eSFQ: eSFQ uses synchronous phase balancing so one junction in each decision-making pair switches during every clock cycle, preventing parasitic circulating currents.This balance removes the need for large ERSFQ feed-line inductances, leaving eSFQ circuits nearly the same area as RSFQ circuits.
- eSFQ: Up to 33–40% more Josephson junctions are required in ERSFQ and eSFQ circuits than in RSFQ circuits.Transitioning existing RSFQ libraries to eSFQ also requires correction and may replace JTLs with shift registers or alternative transmission lines.
6. RSFQ logic family common features
RQL and related SFQ logic address bias-power and synchronization issues through AC or controllable clocking schemes, but introduce frequency, area, and implementation trade-offs.
- Clocking and biasing: ERSFQ clocking can adjust frequency during processing and disable bias voltage for zero-power sleep mode.Logic-cell control adjusts the JTL-generated bias voltage, enabling circuit-level power saving through partitioned biasing.
- Clocking and biasing: RSFQ bias current scales with Josephson-junction count, reaching approximately 100 A for 1 million junctions, while partitioning keeps it below 3 A.Partitioning separates circuits into islands with equal bias current but different bias voltage.
- RQL operation: RQL applies AC power through bias transformers and uses four phase-shifted bias currents to provide directional SFQ propagation.A single AC bias produces periodic flux oscillations; four phases shifted by 0, π/2, π, and 3π/2 provide directionality.
- RQL trade-offs: RQL power supply eliminates cryogenic static bias dissipation and return-current magnetic fields, but high-frequency splitters and transformers increase area and limit miniaturization.In an 8-bit carry-look-ahead adder, power splitters occupied approximately 2.5 times the adder area.
- RQL operation: RQL pipelines provide self-synchronization and an estimated maximum clock frequency of approximately 17 GHz, while pipeline depth trades against clock frequency.Early pulses wait for the next bias phase, limiting accumulated jitter to within one pipeline.
- RQL trade-offs: RQL is limited by multiphase-bias clock skew, approximately 10 GHz practical frequency, absent logic-cell clock control, and RF losses reaching 50% of the power budget.RSFQ circuits routinely operate at 50 GHz, whereas RQL’s multiphase AC design introduces additional high-frequency constraints.
- Energy comparison: Active-mode RQL and ERSFQ power dissipation is roughly similar because their switching counts are comparable for balanced zero-one data.The review reports that only adiabatic switching markedly improves superconducting-circuit energy efficiency.
C. Adiabatic superconductor logic
Adiabatic superconductor logic uses tunable potential landscapes and reversible state transfer to reduce switching energy, but early parametric-quantron processors incurred severe hardware and speed penalties.
- Energy principles: Superconductor logic states in non-adiabatic irreversible circuits are separated by energy barriers of approximately 10^3–10^4 k_BT, below semiconductor barriers of approximately 10^6 k_BT.The minimum barrier approaches Landauer’s limit, Emin = kBT ln 2, where thermal fluctuations threaten state distinguishability.
- Energy principles: Adiabatic switching can reduce the estimated operation energy from EJ ≈ 2 × 10^-19 J to Emin ≈ 4 × 10^-23 J at 4.2 K.The comparison assumes logically irreversible operation and contrasts Josephson-junction switching with the thermodynamic limit.
- Parametric quantron operation: A parametric quantron’s potential energy combines Josephson-junction and magnetic terms controlled by external flux and current.The external parameters tune the potential’s vertex and parabolic-term slope through normalized flux and inductance.
- Parametric quantron operation: Near normalized external flux ϕe ≈ π, changing normalized inductance l switches the potential between single-well and double-well shapes.For l < 1 the potential is single-well; for l > 1, logical states correspond to minima on opposite sides of phase π.
- Reversible transfer: Sequential current pulses transfer logical states through magnetically coupled parametric quantron arrays, with adiabatic pulse shaping enabling reversible operations.The logical state is most pronounced in the cell having the largest normalized inductance at a given moment.
- Reversible transfer: Early reversible parametric-quantron processors were impractical because temporary intermediate-result storage and short-range interactions created severe hardware overhead.An 8-bit 1024-point fast convolver required almost 10^7 parametric quantrons, about 90% used as shift-register elements; the circuits were also slow and parameter-sensitive.
2. Quantum flux parametron based circuits
Quantum-flux-parametron-derived adiabatic circuits offer exceptional energy efficiency and increasing operating speed, while transformer coupling and clock distribution constrain scale and frequency.
- Energy efficiency: AQFP circuits experimentally dissipated approximately 10^-20 J per operation at a 5 GHz clock frequency.Theoretical analysis indicates operation below the thermodynamic limit and possible approach to the quantum limit at 4.2 K.
- Energy efficiency: AQFP implementation of the Collatz algorithm was about seven orders of magnitude more energy efficient than a CMOS FPGA, including cryogenic-cooling power.The comparison used standard manufacturing processes and included the power required for cooling.
- Circuit implementation: AQFP logic can be built from four blocks—buffer, NOT, constant, and branch—together with a latch, and a 10-thousand-gate circuit has been reported.These blocks support adiabatic circuits of arbitrary complexity within the described design framework.
- Scaling constraints: Transformer-based AQFP coupling limits wire length to about 1 mm and practical clock frequency to 5 GHz, with larger circuits expected to develop clock skew.The wire-length constraint reflects the need to maintain sufficient bias flux despite parameter variation.
- Scaling constraints: Increasing activation-current amplitude can exploit AQFP potential periodicity to double or quadruple activation frequency, opening operation at 10 or even 20 GHz.This addresses the relatively low frequency and large latency associated with adiabatic circuits.
3. nSQUID-based circuits
nSQUID-based circuits use magnetic-flux storage and SFQ activation pulses to support adjustable-clock operation with very low estimated dissipation, while circuit design remains constrained by SFQ recycling and integration-density limits.
- nSQUID circuits represent information in SQUID magnetic flux and can receive sequential activation pulses through a common bias bus.
- nSQUID-based transmission lines substitute nSQUIDs for Josephson junctions in RSFQ-like structures and support magnetic or galvanic coupling.
- At 50 MHz, nSQUID logic dissipation was estimated near the thermodynamic limit, ∼2kBT ln 2, after successful tests at 5 GHz.
- SFQ recycling avoids the much larger energy of SFQ creation or annihilation, but closed-loop timing-belt designs impose circuit-design restrictions.
- Low integration density remains the main limitation on superconductor-circuit complexity and performance, motivating miniaturized junctions, kinetic inductance, 3D architectures, and higher-functionality elements.Josephson junction density up to 10^8/cm^2 is described as achievable with planned technological advances.
- Cryogenic memory development is driven by the need for dense RAM, including SQUID, hybrid Josephson-CMOS, JMRAM, and OST-MRAM approaches.
II. MEMORY
The review surveys four competing approaches to cryogenic memory compatible with energy-efficient superconducting electronics.
- The four approaches are SQUID-based memory, hybrid Josephson-CMOS memory, JMRAM, and OST-MRAM.
A. SQUID-based memory
SQUID memory offers very fast access but faces density constraints; hybrid Josephson-CMOS memory improves capacity while adding substantial interface power and timing costs, motivating magnetic-junction alternatives.
- SQUID-based memory stores digital information through the presence or absence of SFQ(s) in a superconducting loop and offers write/read times of a few picoseconds.
- Low integration density limited SQUID-based RAM, leading to hybrid Josephson-CMOS RAM with a 64 kb, 4 K implementation and 400 ps read time.
- Hybrid memory used CMOS cells about three orders of magnitude smaller than SQUID-based counterparts, but required amplification of sub-mV superconducting signals toward CMOS voltage levels.
- Energy-efficient decoders and n-Trons could improve hybrid-memory energy efficiency up to 3 times for 64 kb and up to 12 times for 16 Mb memory.For 16 Mb memory, estimated read access time is 0.78 ns.
- Practical low-temperature RAM targets include element dimensions below 100 nm, 10^-18 J writes in roughly 50–100 ps, and 10^-19 J reads in roughly 5 ps.
- Magnetic Josephson junctions are proposed to reduce memory-cell size, but their generally low characteristic frequency slows readout and complicates SFQ integration.
2. MJJ valve with in-plane heterogeneity of the weak-link area
In-plane heterogeneous MJJ valves use differing CPR regions and magnetization states for memory operation, while addressing slow writing and integration challenges through phase-based and spin-transfer approaches.
- 2. MJJ valve with in-plane heterogeneity of the weak-link area: Heterogeneous MJJ valves separate the weak-link region into parts with different CPRs, including conventional and π-shifted regions, forming a nanoSQUID-like structure.
- 2. MJJ valve with in-plane heterogeneity of the weak-link area: For SF-NFS-based valves, perpendicular magnetization can compensate the Josephson phase gradient and produce high critical current, whereas 90° rotation lowers it; residual flux must be comparable to Φ0.
- 2. MJJ valve with in-plane heterogeneity of the weak-link area: MJJ memory commonly has slow writes because magnetization reversal takes longer than SIS switching, while vortex-based fast writing dissipates approximately 10^-18 J during vortex annihilation.
- 2. MJJ valve with in-plane heterogeneity of the weak-link area: Bistable Josephson potentials with ground-state phases ±ϕ can support picosecond write and read operations, but implementing the ϕ-state requires a relatively large heterogeneous weak-link structure.
- D. OST-MRAM: OST memory uses orthogonal spin-transfer torque to rotate a free-layer magnetization, and sub-ns switching has been demonstrated with approximately mA current pulses.
- D. OST-MRAM: OST devices eliminate magnetic-field control lines and enable fast reversal, but low magnetoresistance and possible magnetization over-rotation still impede practical application.
- D. OST-MRAM: Applying spin-transfer torque to MJJ valves remains an open research direction because coupled superconducting and magnetic dynamics require further evaluation.
E. Discussion
The discussion identifies low integration density and cryogenic RAM as major barriers, while outlining material, fabrication, architectural, and cross-platform directions for advancing superconducting computing.
- Memory challenges: Cryogenic RAM remains a major obstacle, with four approaches considered but no clear winner.Further progress depends on operation principles combining superconductivity and magnetism.
- Memory challenges: Memory development is constrained by thin magnetic layers, exponential thickness sensitivity, and interfacial roughness.Higher-accuracy thin-film fabrication is identified as a way to address these constraints.
- Memory challenges: Address lines likely dominate memory-matrix area, delay, and dissipated power, making interconnections and cell architecture important optimization targets.
- Capabilities and directions: Superconducting logic supports fast operation at approximately 5–50 GHz and energy-efficient operation at 10^-19–10^-20 J per bit.The discussion covers both non-adiabatic and adiabatic regimes.
- Capabilities and directions: Combining logic schemes and superconducting with semiconductor technologies could support unconventional computational paradigms and cryogenic cross-platform supercomputers.Proposed paradigms include cellular automata, artificial neural networks, adiabatic, reversible, and quantum computing.
- Technology limitations: Low integration density limits the functional complexity of superconducting devices and requires miniaturization, modernized cell libraries, and novel devices.Suggested device directions include nanowires and magnetic Josephson junctions.