Source-linked AI summary
Neural-like computing with populations of superparamagnetic basis functions
Alice Mizrahi, Tifenn Hirtzlin, Akio Fukushima, Hitoshi Kubota, Shinji Yuasa, Julie Grollier, Damien Querlioz
TL;DR
The paper addresses the challenge of implementing fault-tolerant population coding with noisy, variable nanodevices. It uses superparamagnetic tunnel junctions as stochastic neuron-like elements, assembles their tuning curves into basis functions, and connects populations with CMOS and magnetic memory. Experimentally, nine junctions generate nonlinear functions and cursive letters, while modeled populations learn nonlinear transformations despite device variability in a low-area, low-energy architecture.
Problem
Nanodevices need nonlinear, differently tuned responses to realize population coding, but this had not been demonstrated experimentally with nanodevices.
Method
The paper assembles superparamagnetic tunnel junctions with shifted current-dependent tuning curves into weighted basis-function populations and hybrid CMOS-magnetic systems.
Results
Nine experimental junctions implement a basis set that reconstructs nonlinear functions and generates six cursive letters, while modeled interconnected populations learn nonlinear transformations with device variability.
Takeaways & Limitations
Superparamagnetic tunnel junction populations provide a hardware substrate for stochastic, variability-resilient population coding and cascaded nonlinear computation.
Abstract
from arXiv · showhide
In neuroscience, population coding theory demonstrates that neural assemblies can achieve fault-tolerant information processing. Mapped to nanoelectronics, this strategy could allow for reliable computing with scaled-down, noisy, imperfect devices. Doing so requires that the population components form a set of basis functions in terms of their response functions to inputs, offering a physical substrate for calculating. For this purpose, the responses of the nanodevices should be non-linear, and each tuned to different values of the input. These strong requirements have prevented a demonstration of population coding with nanodevices. Here, we show that nanoscale magnetic tunnel junctions can be assembled to meet these requirements. We demonstrate experimentally that a population of nine junctions can implement a basis set of functions, providing the data to achieve, for example, the generation of cursive letters. We design hybrid magnetic-CMOS systems based on interlinked populations of junctions and show that they can learn to realize non-linear variability-resilient transformations with a low imprint area and low power.
Tuning curve of a superparamagnetic tunnel junction
Superparamagnetic tunnel junctions provide neuron-like, current-dependent tuning curves that can be assembled into basis functions for computation. Populations and interconnected systems use these responses to learn nonlinear transformations while tolerating variability and supporting low-area, low-energy implementations.
- Device operation: Thermal fluctuations in nanoscale magnetic tunnel junctions produce stochastic switching between parallel and antiparallel states.The reduced energy barrier allows sustained oscillations between the two magnetic configurations.
- Device operation: Spin-transfer torque makes the switching rate depend nonlinearly on applied current, producing a neuron-like tuning curve.Positive and negative currents stabilize opposite magnetic states and reduce switching rates relative to near-zero current.
- Basis functions: The junction response approximates a Gaussian tuning curve over a narrow current range, enabling a basis set with shifted peak positions.The reported sensing range is approximately ±50 µA around zero current.
- Experimental computation: Nine experimentally measured junctions reconstruct an altimeter function and generate six cursive letters through weighted combinations of their tuning curves.The letter examples are w, i, n, r, u, and m.
- Learning and transformations: Interconnected populations learn identity and nonlinear transformations, with the gripper task reaching a mean error below 2.5% after 3,000 learning steps.The learned transformations include linear scaling, square, inverse, sine, and cascaded sin^2(x).
- System implementation: The proposed hybrid implementation combines magnetic junctions, CMOS, and ST-MRAM for low-area, low-energy stochastic computing.A 128-input, 128-output design occupies 12,000 µm² and consumes 23 nJ during learning versus 7.4 nJ afterward; precision trades off against observation time and energy.
METHODS
The methods characterize experimentally fabricated magnetic tunnel junctions and fit analytical models to their measured responses. Experimental data and target functions are then used to determine weights and parameterize simulations.
- Fabrication: 60 × 120 nm^2 nanopillars were fabricated from sputtered multilayer magnetic tunnel junction stacks after annealing at 300°C under a 1-T magnetic field.The stack included IrMn, CoFe, Ru, CoFeB, MgO, and capping layers.
- Characterization: The measured junction curves were shifted along the current axis after measurements under a field canceling the synthetic-antiferromagnet stray field.The curves were initially centered on zero voltage.
- Model fitting: Equation 2 was fit to experimental data by choosing ΔE and I_c separately for each junction.The nine junctions used distinct fitted energy-barrier and critical-current parameters.
- Variability: Parameter variability was attributed to polycrystalline free-ferromagnet structure and partial rather than full-layer magnetization reversal.This mechanism explains both instability and strong device-to-device parameter variation.
- Weight determination: Weights for the experimental basis-set reconstruction were obtained analytically from the measured target values and rate matrix, whereas later figures used learning without matrix inversion.The reconstruction uses H = w*R and w = H*R^-1.
- Simulation target: The simulated target function used a measured-point height model with α = 5.255 and A/T0 = 2.26 × 10^-5.The implementation expressed height as a function of direct current.
- Simulation parameters: Simulation parameters were drawn from experimentally extracted distributions: ΔE centered at 13.78 kBT with 9.65 kBT span and V_c with mean 0.142 V and standard deviation 0.037 V.These distributions were used to represent junction variability and natural switching behavior.
Simulations of a population of junctions
The simulations model junction populations as stochastic two-state devices driven by shared voltages and connect input and output populations through computed voltages and adjustable weights. They evaluate sensing, learning, and orientation error over a defined voltage range.
- Population model: Each junction is modeled as a two-state Poisson process whose escape rates are modified by the applied stimulus.Voltage control allows one common stimulus to be applied to all junctions.
- Input encoding: For each input junction, effective voltage is V_eff = V − V_0, where V is the common stimulus and V_0 is that junction’s tuned voltage.V_c is the critical voltage in the switching-probability model.
- Numerical procedure: At every time step, switching probabilities are evaluated and randomized state changes are simulated; frequencies are computed after 100 steps of dt = 439 μs.The procedure is repeated independently for each junction.
- Population interconnection: Output-junction voltages are computed by inverting the junction-rate model so their rates satisfy the desired transformation equation.The resulting output population is then simulated using the same stochastic procedure.
- Input range: The input population senses voltages from −0.15 V to +0.15 V, encoding possible object orientations.Shifting individual junction rates enables sensing of different ranges for coordinate transformations.
- Learning: Learning updates increase or decrease input-to-output weights according to whether the corresponding connections should be strengthened or weakened.The supplied simulation description specifies separate update rules for the two cases.
- Learning: The learning rate was set to α = 0.001 because lower values slow learning while higher values accelerate it but limit performance.R_0 denotes the natural junction rate.
- Evaluation: Error is the absolute orientation difference between target and gripper output, expressed as a percentage of the −0.15 V to +0.15 V orientation range and averaged over 50 trials.This metric quantifies orientation performance in the simulated task.
1-dimension coordinate transformations
The one-dimensional transformation task replaces the object orientation with a transformed target and measures the gripper’s distance from that expected value. Identity and doubling use the same input voltage range.
- Transformation evaluation: The transformation task substitutes T(Z) for the object orientation and computes gripper-target distance as the absolute difference from the expected T(Z).The distance is reported as a percentage of the range of possible expected values.
- Transformation cases: For Identity, T(Z) = Z, and Double, T(Z) = 2Z, the stimulus range is −0.15 V to +0.15 V.These two transformations provide the specified one-dimensional test cases.
2-dimensional coordinate transformation
The system learns a two-dimensional coordinate transformation by connecting populations encoding R and φ to output populations encoding x and y. Its hybrid junction-CMOS implementation uses fixed-point computation and is estimated for low area and energy consumption.
- x = R cos(φ π/0.6) and y = R sin(φ π/0.6) define the coordinate transformations.
- Four junction populations encode the input coordinates R and φ and the output coordinates x and y, each spanning 0 to 0.3 V.
- The R and φ populations are concatenated, then weight matrices Wx and Wy connect them to the x and y output populations.
- The implementation also evaluates a sequential two-step square-of-sine transformation using three populations and two trained weight matrices.
- P = 3.2 µW is the estimated total power for the modeled 100-junction system, combining shifting and stimulus power.The reported components are Pshift = 0.8 µW and Pstim = 2.4 µW.
- The CMOS system is designed with standard 28 nm integrated-circuit tools and optimized for low area and low energy rather than high speed.
SECTION 1: USING SPIN-ORBIT TORQUES TO SHIFT THE JUNCTIONS
Spin-orbit torque shifts the tuning curves of individual superparamagnetic junctions, allowing a population to contain junctions tuned to different voltages. The shift is controlled through the geometry of a heavy-metal underlayer.
- A current injected into a heavy-metal underlayer modifies the free-layer magnetization and shifts the junction tuning curve.
- The frequency F(V,w) depends on the underlayer width w and the junction parameters, while the common stimulus is VSTT.
- Choosing different heavy-metal underlayer shapes produces different effective biases and tunes junctions to different voltages.
SECTION 2: ROBUSTNESS TO VARIABILITY
The system is reported to be robust to variability in junction critical voltage because the tuning-curve width scales with that voltage. Moderate variability can improve precision, whereas excessive variability worsens the match to theoretical curves.
- The effect of critical-voltage variability is evaluated using populations of 100 junctions and 3,000 learning steps, averaging each point over 5 trials.Error bars represent the corresponding standard deviation.
- Random critical-voltage variations do not affect coding precision because tuning-curve width is proportional to the critical voltage.
- Uniform energy-barrier variability can raise the average frequency above the theoretical frequency F0 and increase precision.
- When energy-barrier variability is too high, mismatch between expected and observed tuning curves makes precision worse than without variability.
SECTION 3: RESILIENCE TO THE LOSS OF NEURONS
The study evaluates how population coding responds to junction loss and finds that the system retains useful performance and relearns rapidly after failures.
- SECTION 3: RESILIENCE TO THE LOSS OF NEURONS: The simulations model neuron loss by setting randomly selected junction rates to zero after training.The system uses input and output populations of 100 junctions, with loss applied after 3000 learning steps.
- SECTION 3: RESILIENCE TO THE LOSS OF NEURONS: Loss increases the distance gripper-target, but performance remains better than that of an untrained network even without relearning.This indicates that the surviving population continues contributing useful computation after failures.
- SECTION 3: RESILIENCE TO THE LOSS OF NEURONS: Final performance after relearning matches performance from training a system that experienced the same loss before learning.The comparison holds across the tested loss levels.
- SECTION 3: RESILIENCE TO THE LOSS OF NEURONS: Several hundred relearning steps are sufficient after loss, compared with several thousand steps for initial learning.The authors interpret this shorter recovery as rapid adaptation to drastic changes.
- SECTION 3: RESILIENCE TO THE LOSS OF NEURONS: The authors conclude that the computing system is resilient to faulty superparamagnetic tunnel junctions.The section links this conclusion to the measured recovery behavior after neuron loss.
SECTION 5: DATA PATH OF THE FULL SYSTEM.
The full system combines superparamagnetic junctions, ST-MRAM weights, CMOS circuitry, and finite-state control to count switching events, compute weighted outputs, and update weights during learning.
- DATA PATH: The datapath associates superparamagnetic junctions with ST-MRAM weights and CMOS circuitry, using fixed-point integer arithmetic.The schematic also identifies RAM, finite-state-machine, write-enable, and address signals.
- DATA PATH: Before operation, the system programs ST-MRAM weights with random values.These values provide the starting state before computation or learning proceeds.
- DATA PATH: In state S0, output counters count switching events from each superparamagnetic tunnel junction.The counters convert junction activity into values used by later datapath stages.
- DATA PATH: When enabled, states S1–S3 multiply input counts by stored weights, accumulate the results, and store output values in registers.Address increment is automatic during the sequential computation of the output values.
- DATA PATH: During learning, states S4–S6 compare outputs with input addresses and update ST-MRAM weights according to the learning rule.The update uses additions and multiplications involving the system parameters.
- READOUT: The junction readout compares junction and reference-resistor voltages with a low-power CMOS comparator before digital counting.The design cannot detect multiple switching events within one clock cycle, although system-level simulations report no application-level impact.
SECTION 7: COMPARISON WITH PURELY CMOS BASED OPTIONS.
The paper compares its hybrid magnetic-CMOS design with fully CMOS alternatives, emphasizing lower reported area and spike-generation energy while noting trade-offs in analog approaches.
- COMPARISON: The proposed architecture converts analog inputs into digital spiking populations and combines them with digital circuitry and local memory for learning.The paper compares this hybrid design with several entirely CMOS options.
- AREA AND ENERGY: The reference CMOS spiking-neuron design would occupy 1.3mm², whereas the proposed whole circuit occupies 0.12mm².The comparison concerns the area of the neuron design versus the full proposed circuit.
- AREA AND ENERGY: Generating one population’s spikes would consume 330nJ in the reference design versus 0.22nJ in the proposed approach.These values are reported for the same comparison of spike-generation energy.
- CMOS ALTERNATIVES: An analog non-spiking CMOS option is estimated at 1,280µm² and 200 pJ for a 10 µs run, but its outputs require conversion before digital processing.The required ADC can become dominant in area and energy when conversion is performed per neuron.
- CMOS ALTERNATIVES: Analog computation with one output ADC remains limited in scalability and makes memory implementation harder.The paper nevertheless identifies it as attractive when superparamagnetic tunnel junctions are unavailable.
- CMOS ALTERNATIVES: A fully digital option using one input ADC is highly scalable, but requires additional digital circuitry to compute population values.The paper reports one ADC at 20nJ and 0.2mm² for this option.
SECTION 8: ADAPTATION OF THE SYSTEM TO MULTIPLE INPUTS.
The multi-input architecture connects each output population to both input populations to implement a transformation from polar to Cartesian coordinates.
- MULTIPLE INPUTS: Each output population, X and Y, is linked to both input populations, R and φ.This two-input architecture implements the transformation from polar to Cartesian coordinates.