Source-linked AI summary
Vowel recognition with four coupled spin-torque nano-oscillators
Miguel Romera, Philippe Talatchian, Sumito Tsunegi, Flavio Abreu Araujo, Vincent Cros, Paolo Bortolotti, Juan Trastoy, Kay Yakushiji, Akio Fukushima, Hitoshi Kubota, Shinji Yuasa, Maxence Ernoult, Damir Vodenicarevic, Tifenn Hirtzlin, Nicolas Locatelli, Damien Querlioz, Julie Grollier
TL;DR
Training neural networks built from dynamical nanodevices requires precise control of coupled oscillations. This paper trains four spin-torque nano-oscillators using automatic real-time frequency tuning to recognize spoken vowels, achieving recognition rates up to 89% on training data and 88% on testing data. The results show that small hardware networks using oscillations and synchronization can perform non-trivial pattern classification.
Problem
Training neural networks composed of dynamical nanodevices requires finely controlling and tuning their coupled oscillations.
Method
The authors train a hardware network of four spin-torque nano-oscillators by tuning their frequencies with an automatic real-time learning rule.
Results
89% training and 88% testing recognition rates are achieved for spoken-vowel classification.
Takeaways & Limitations
Small hardware neural networks can perform non-trivial pattern classification using non-linear dynamical features including oscillations and synchronization.
Takeaways & Limitations
The dynamical neural networks will need to be scaled up for challenging classification problems on software-benchmarked databases.
Abstract
from arXiv · showhide
Substantial evidence indicates that the brain uses principles of non-linear dynamics in neural processes, providing inspiration for computing with nanoelectronic devices. However, training neural networks composed of dynamical nanodevices requires finely controlling and tuning their coupled oscillations. In this work, we show that the outstanding tunability of spintronic nano-oscillators can solve this challenge. We successfully train a hardware network of four spin-torque nano-oscillators to recognize spoken vowels by tuning their frequencies according to an automatic real-time learning rule. We show that the high experimental recognition rates stem from the high frequency tunability of the oscillators and their mutual coupling. Our results demonstrate that non-trivial pattern classification tasks can be achieved with small hardware neural networks by endowing them with non-linear dynamical features: here, oscillations and synchronization. This demonstration is a milestone for spintronics-based neuromorphic computing.
through standard arithmetic operations. By contrast, a prominent branch of neuroinspired
A neuroinspired computing approach assigns dynamical functionality to network components and exploits emergent synchronization to solve complex problems with small networks. This approach is well suited to compact, energy-efficient nanoelectronic hardware using nonlinear auto-oscillators and dynamical couplings.
- Dynamical neuroinspired computing: Neuroinspired networks can assign dynamical functionality, such as oscillations, to each component and use emergent synchronization to compute complex problems with small networks.These dynamical features are proposed as alternatives to computing through standard arithmetic operations.
- Hardware implementation: This approach is especially attractive for hardware because nanoelectronic devices can provide compact, energy-efficient nonlinear auto-oscillators that mimic biological neurons’ periodic spiking activity.The devices are described as hardware components with dynamical functionality.
- Coupled oscillator communication: Dynamical couplings between oscillators can mediate synaptic communication between neurons.The passage introduces dynamical coupling as a mechanism for communication within oscillator-based neural networks.
challenge towards implementing these models with nano-devices is to achieve learning, which
Learning in nano-device models requires finely controlling coupled oscillations, while spintronic nano-oscillators provide tunable frequencies that enable this control. The study trains four such oscillators to recognize spoken vowels using automatic real-time frequency tuning.
- Challenge: Learning requires finely controlling and tuning coupled oscillations in dynamical nano-devices.Nanodevice dynamics can also be difficult to control and susceptible to noise and variability.
- Approach: Spintronic nano-oscillators address this challenge through wide and accurate electrical-current and magnetic-field control of their frequencies.
- Demonstration: A hardware network of four spin-torque nano-oscillators was successfully trained to recognize spoken vowels.
- Learning rule: The oscillators learned by tuning their frequencies according to an automatic real-time learning rule.
- Result: High experimental recognition rates stemmed from the oscillators’ outstanding frequency tunability.
these oscillators to synchronize. Our results demonstrate that non-trivial pattern classification tasks
The results show that small hardware neural networks can perform non-trivial pattern classification using non-linear dynamical features, specifically oscillations and synchronization, with real-time learning.
- Small hardware neural networks can achieve non-trivial pattern classification tasks.
- The networks are endowed with non-linear dynamical features to support classification.
- The demonstrated dynamical features are oscillations and synchronization, alongside real-time learning.
array of four spin-torque nano-oscillators is a milestone for spintronics-based neuromorphic
A hardware network of four mutually coupled spin-torque nano-oscillators was trained in real time to recognize spoken vowels by tuning oscillator frequencies. It achieved high recognition rates because frequency tunability and mutual coupling produce useful synchronization dynamics.
- Four spin-torque nano-oscillators were electrically interconnected so each oscillator’s microwave emission influenced the frequencies and dynamics of the others.The symmetric interconnections used millimeter-long wires and microwave spin-torques.
- The network classified spoken vowels by mapping input formant frequencies to different oscillator synchronization configurations.Inputs were encoded in the frequencies fA and fB of two fixed-amplitude microwave signals.
- After 48 training steps, the oscillator dc currents and frequencies stopped evolving, while recognition rates reached up to 89% on training data and 88% on testing data.The plateau indicated that the automatic real-time learning process had reached an optimum.
- Recognition increased linearly with oscillator locking ranges because larger synchronization regions encompassed more points in the vowel clouds.Mutual coupling enhanced the locking ranges and thereby increased recognition rates.
- 89% experimental recognition approached the 94% maximum obtained with the same neural network using ideal, noiseless oscillators.The reported high performance was attributed to large experimental locking ranges arising from tunability, coupling, and low noise.
- The experimental oscillatory network exceeded 84% recognition with only 30 trained parameters, demonstrating efficient pattern recognition through coupled dynamics and synchronization.The network used 350 vowel presentations and combined complex coupled dynamical features with collective synchronization to inputs.
Methods · A. Samples
The samples were circular magnetic tunnel junctions fabricated from a specified multilayer stack and processed by annealing, etching, and electron-beam lithography. Their vortex-based oscillators exhibited tunable gyration frequencies from 150 MHz to 450 MHz.
- A. Samples: The magnetic tunnel junction films used a buffer/PtMn/CoFe/Ru/CoFeB/CoFe/MgO/FeB/MgO/Ta/Ru multilayer stack with thicknesses specified in nanometers.The stack included PtMn(15), Co71Fe29(2.5), Ru(0.9), Co60Fe20B20(1.6), Co70Fe30(0.8), MgO(1), Fe80B20(6), MgO(1), Ta(8), and Ru(7).
- A. Samples: The films were prepared by ultra-high vacuum magnetron sputtering and annealed at 360 °C for 1 h.These processing conditions followed deposition of the multilayer films.
- A. Samples: The junctions were patterned using Ar ion etching and electron-beam lithography, producing samples with resistance close to 40 Ω.The reported resistance characterizes the patterned samples.
- A. Samples: The samples had a magneto-resistance ratio of about 100 % at room temperature.The FeB layer formed a single magnetic vortex ground state for the device dimensions used.
- A. Samples: The vortex core was about 12 nm in diameter at remanence, with magnetization spiraling out of plane.This vortex structure provided the magnetic state underlying oscillator operation.
- A. Samples: Under dc current injection and spin-transfer torques, the vortex core steadily gyrated around the dot center.The resulting oscillators operated across a frequency range of 150 MHz to 450 MHz.
B. Database and inputs
The study classifies seven spoken vowels using formant data from 37 female speakers in the Hillenbrand database. Three formants sampled at four times are transformed into two oscillator-range input frequencies for the coupled nano-oscillator network.
- Database: The network classifies seven spoken vowels using formants obtained from a subset of the Hillenbrand database.Spoken vowels are characterized by formant frequencies.
- Inputs: Each vowel is represented by F1, F2, and F3 sampled at four duration points, yielding 12 parameters.The sampling times are steady state and 20%, 50%, and 80% of vowel duration.
- Database: 37 female speakers form the resulting database after incomplete formant utterances are removed and speaker counts are equalized across vowels.Entries with unmeasured or unresolved formants were excluded rather than retained as zero-valued parameters.
- Inputs: Two linear combinations of the formants produce characteristic frequencies fA and fB within the oscillators’ 325 MHz–380 MHz operating range.The coefficients are selected using a synchronization-map calibration and least-square fitting, after which fA and fB drive the network as microwave inputs.
C. Experimental set-up
The experiment uses four electrically series-connected vortex spin-torque nano-oscillators, independently controlled by dc currents and driven by two external microwave inputs. Network outputs are analyzed through oscillator synchronization patterns to classify vowels and automatically adjust currents during training.
- Hardware network: Four vortex nano-oscillators are connected in series and coupled through emitted microwave currents, while their separation prevents dipolar-field coupling.The setup schematic shows the four coupled oscillators; millimeter-long wires provide electrical coupling, and the oscillators are too far apart for magnetic dipolar coupling.
- Hardware network: Four independent dc sources control the oscillator currents through cumulative current paths, enabling separate tuning of each device.The actual currents are ISTO1=IDC1, ISTO2=IDC2+IDC1, ISTO3=IDC3+IDC2+IDC1 and ISTO4=IDC4+IDC3+IDC2+IDC1.
- Input and measurement: Two microwave signals with frequencies fA and fB and power P = -9 dBm are injected through a strip line as inputs to the oscillator network.The strip line is 2.5 µm wide, positioned 370 nm above the pillar, and produces input microwave fields with amplitude 0.1 mT.
- Input and measurement: A spectrum analyzer records microwave emissions, and the output oscillator frequencies are used to classify spoken vowels according to input-dependent dynamics.The analysis extracts four oscillator frequencies in the presence of microwave inputs, whose output depends on the microwave input frequencies.
- Recognition and learning: Detected frequencies within ± 0.5 MHz of an external signal count as synchronized, and the resulting pattern is compared with the vowel’s assigned pattern.The real-time program checks whether the applied vowel was properly recognized by comparing detected synchronization states with the initially assigned pattern.
- Recognition and learning: During training, misclassification triggers an online algorithm that calculates dc-current changes to reduce recognition error and automatically applies them to the setup.The modified-current information is sent back to the experimental system, where the dc currents are automatically changed.
D. Real-time learning algorithm · E. Cross-validation procedure · F. Comparison of spin-torque nano-oscillators to CMOS oscillators
The network learns seven vowel classes by adjusting oscillator currents to align measured synchronization configurations with assigned patterns. Performance is evaluated through five-fold cross-validation, alongside a feature comparison with CMOS oscillators.
- D. Real-time learning algorithm: The supervised procedure recognizes seven spoken English vowels by assigning each class a synchronization pattern.The classes are “AE”, “AH”, “AW”, “ER”, “IH”, “IY” and “UW”.
- D. Real-time learning algorithm: For each vowel, a four-component frequency difference vector measures the distance between the applied input and its assigned synchronization region.The four components correspond to the four oscillators.
- D. Real-time learning algorithm: The automatic rule iteratively modifies oscillator dc currents according to the summed frequency-difference vectors, displacing synchronization patterns toward an optimal configuration.The learning rate is μ = 0.1 mA, and each oscillator current changes only by ±μ per step.
- D. Real-time learning algorithm: 89 % recognition was reached after step 48, after which the rate saturated; training used a maximum of N = 87 steps.N = 87 corresponds to applying three times each of the 29 training-database datapoints.
- E. Cross-validation procedure: Training used 80% of the vowel database, while testing used the remaining 20% of data points.The procedure was repeated five times using distinct testing samples.
- E. Cross-validation procedure: Five-fold cross-validation used successive 20% data-point quintiles for testing, and the final recognition rate averaged the five testing rates.The same cross-validation procedure was applied to all neural networks.
- F. Comparison of spin-torque nano-oscillators to CMOS oscillators: Extended Data Table 2 compares CMOS and spin-torque nano-oscillators for neuromorphic computing.The vortex spin-torque oscillators are the magnetic tunnel junctions used in this study, while 10 nm spin-torque oscillators are memory-cell junctions.
G. Comparison with a multilayer perceptron
A standard multilayer perceptron was benchmarked on the same vowel database, using 12 formants as inputs and seven vowel-class outputs. Its recognition rate was evaluated across hidden-layer sizes and trained-parameter counts after cross-validation.
- Results: Results were reported as recognition rate after cross-validation versus the number of trained parameters.The perceptron served as a benchmark for the experimental oscillatory network on the same vowel database.
- Architecture: The perceptron used 12 vowel formants as inputs and seven outputs, one for each vowel class.Formants were rescaled between -1 and 1 before entering the first layer.
- Architecture: Hidden-neuron counts were varied from 1 to 20 to evaluate recognition rate as a function of trained parameters.The output class was selected from the largest softmax output.
- Results: ReLU activation functions performed worse than tanh on this vowel-recognition task.The hidden layer used tanh activations, while the output layer used softmax.
- Training: The network was trained by backpropagation with gradient descent over negative log-likelihood, using randomly presented samples and tuned learning rate.Each iteration comprised a forward pass, gradient evaluation, and weight update.
H. Comparison with recurrent neural networks
The study compares a perceptron, multilayer perceptron, RNN, and LSTM with four hidden units on the same vowel database. Their recognition performance is evaluated by cross-validation success as a function of the number of learned parameters.
- Architectures: Four neural-network architectures were evaluated on the same vowel database: a perceptron, multilayer perceptron, RNN, and LSTM with four hidden units.The architectures and their schematics are reported in Extended Data Fig. 2.
- Training procedure: Formants were presented sequentially, with softmax output and tanh elsewhere, and the vowel class was selected from the maximum activation in a seven-class one-hot encoding.Samples were selected and presented randomly during training.
- Optimization: Learning rates were tuned for each architecture, while initial weights and biases were sampled from a Gaussian with mean 0 and variance 0.01.No gradient inertia or learning-rate adaptation was used.
- Optimization: 500000 and 1000000 training iterations were used for the LSTM and RNN, respectively, with learning rates of 0.01 and 0.0005.These training durations were used to ensure convergence.
- Results: Extended Data Fig. 2a reports cross-validation success as a function of the number of learned parameters for the compared networks.Test and training rates were averaged over the last 5000 iterations to obtain reliable trial estimates.
I. Synchronization detection through oscillator rectified voltages … 4 Evaluation of the injection locking range normalized by the frequency difference between
The study develops synchronization detection through spin-diode rectified voltages, models four coupled oscillators, and evaluates how tunability, mutual coupling, and locking-range normalization shape recognition performance. Simulations identify synchronization conditions from frequency differences and compare oscillator-based detection energy with CMOS implementation estimates.
- I. Synchronization detection through oscillator rectified voltages: Synchronization is detected when a rectified voltage appears, coinciding with the oscillator’s injection-locking range.The voltage arises from the spin-diode effect and is proportional to the external microwave current fraction flowing through the oscillator.
- I. Synchronization detection through oscillator rectified voltages: The proposed CMOS differential circuit outputs a binary high or low voltage indicating whether an oscillator is synchronized to its input.It compares each oscillator voltage with a similarly polarized reference resistor using two differential stages and a gain stage.
- 1. Model description: The modeled network contains four van der Pol oscillators with a 2% natural-frequency mismatch and two microwave inputs.The model captures essential coupled dynamics of spin-torque nano-oscillators and includes normalized mutual coupling.
- 2. Recognition performances: Recognition is optimized when the oscillators have similar free-running frequency differences and similar injection-locking-range widths.For each oscillator-parameter set, the formant transformation is adapted and the learning process is simulated to obtain the optimum recognition rate.
- 2. Recognition performances: The tunability study varies the normalized nonlinear frequency-shift coefficient N0 from 0.00 to 0.26 in steps of 0.02 without mutual coupling.The coupling study fixes N0 = 0.08, corresponding to a locking range/frequency difference of 0.58, and varies normalized mutual coupling ε; the experimental ε is 1.6.
- 4 Evaluation of the injection locking range normalized by the frequency difference between: The locking-range-to-frequency-difference ratio is computed from the average locking range of four oscillators divided by adjacent natural-frequency differences.The adjacent differences are δ12 = |ω1 − ω2|, δ23 = |ω2 − ω3|, and δ34 = |ω3 − ω4|.