Source-linked AI summary

Large-Scale Neuromorphic Spiking Array Processors: A quest to mimic the brain

Chetan Singh Thakur, Jamal Molin, Gert Cauwenberghs, Giacomo Indiveri, Kundan Kumar, Ning Qiao, Johannes Schemmel, Runchun Wang, Elisabetta Chicca, Jennifer Olson Hasler, Jae-sun Seo, Shimeng Yu, Yu Cao, André van Schaik, Ralph Etienne-Cummings

arXiv:1805.08932v1cs.NE

TL;DR

Neuromorphic engineering seeks brain-inspired, efficient computation to address large-scale neural simulation and growing data-processing demands. This review compares major neuromorphic spiking emulators, their architectures, advantages, drawbacks, and capabilities, finding that systems offer complementary trade-offs while advancing scalable, energy-efficient neural simulation.

  • Problem

    Large-scale brain simulation and rising data-processing demands exceed the real-time efficiency of conventional digital systems, motivating brain-inspired hardware approaches.

  • Method

    The review compares significant neuromorphic spiking emulators and the architectures, approaches, advantages, drawbacks, and modeling capabilities they provide.

  • Results

    The surveyed systems provide distinct capabilities and trade-offs, including accelerated model execution, scalable reconfiguration, parallel co-localized computation, and biologically oriented or digitally robust neuron implementations.

  • Takeaways & Limitations

    Together, these complementary systems move neuromorphic engineering toward compact, energy-efficient, dense neural simulators and real-time processing devices.

  • Takeaways & Limitations

    Energy efficiency remains primarily limited by external memory access for lookup of fully reconfigurable, sparse synaptic connectivity.

Abstract

from arXiv · show

Neuromorphic engineering (NE) encompasses a diverse range of approaches to information processing that are inspired by neurobiological systems, and this feature distinguishes neuromorphic systems from conventional computing systems. The brain has evolved over billions of years to solve difficult engineering problems by using efficient, parallel, low-power computation. The goal of NE is to design systems capable of brain-like computation. Numerous large-scale neuromorphic projects have emerged recently. This interdisciplinary field was listed among the top 10 technology breakthroughs of 2014 by the MIT Technology Review and among the top 10 emerging technologies of 2015 by the World Economic Forum. NE has two-way goals: one, a scientific goal to understand the computational properties of biological neural systems by using models implemented in integrated circuits (ICs); second, an engineering goal to exploit the known properties of biological systems to design and implement efficient devices for engineering applications. Building hardware neural emulators can be extremely useful for simulating large-scale neural models to explain how intelligent behavior arises in the brain. The principle advantages of neuromorphic emulators are that they are highly energy efficient, parallel and distributed, and require a small silicon area. Thus, compared to conventional CPUs, these neuromorphic emulators are beneficial in many engineering applications such as for the porting of deep learning algorithms for various recognitions tasks. In this review article, we describe some of the most significant neuromorphic spiking emulators, compare the different architectures and approaches used by them, illustrate their advantages and drawbacks, and highlight the capabilities that each can deliver to neural modelers.

1 Introduction

Neuromorphic engineering applies brain-inspired computation to address the energy and scalability limits of conventional large-scale neural simulation. The review surveys neuromorphic processors, their architectures, strengths, drawbacks, and applications.

  • Mouse-scale cortical simulation on a personal computer uses 40,000 times more power and runs 9000 times slower than a mouse brain.Human-scale simulation is projected to require an exascale supercomputer and 0.5 GW of power.
  • Neuromorphic computing offers a brain-inspired alternative for handling increasing data-processing requirements.The approach is motivated by differences between analog neural computation and conventional digital computing.
  • Silicon neurons emulate neuronal and synaptic electrophysiology directly in hardware, enabling energy-efficient, real-time large-scale neural emulations.Network speed can be independent of neuron count or coupling.
  • Neuromorphic processors are being developed for real-time recognition and emerging applications including autonomous cars, drones, and brain-machine interfaces.Commercial examples include touchpad, biometric, color-imaging, artificial-retina, and dynamic-vision technologies.
  • Parallel, redundant, and co-localized memory-computation architectures support efficient pattern recognition while addressing the von Neumann memory bottleneck.Mixed-signal processors can also reduce silicon-area usage relative to pure digital solutions.
  • The review compares neural processors spanning current-mode sub-threshold to voltage-mode switched-capacitor designs and discusses their strengths and applications.

2 Integrate-and-Fire Array Transceiver (IFAT)

IFAT neuromorphic arrays use event-driven, reconfigurable architectures to implement adaptive integrate-and-fire neurons, with MNIFAT demonstrating compact, low-power operation and controlled mismatch.

  • IFAT architecture: IFAT combines mixed-mode VLSI neurons, reconfigurable weighted synapses, lookup-table connectivity, and asynchronous AER event communication.AER addresses identify receiving neurons, while the architecture supports event-based input and output routing.
  • MNIFAT design: MNIFAT integrates 2040 Mihalas-Niebur neurons, each capable of operating as two independent leaky integrate-and-fire neurons, yielding 4080 I&F neurons.The design uses adaptive thresholds and was implemented in 0.5 µm CMOS at a 5 V nominal supply.
  • Neuron model: The modified Mihalas-Niebur circuit models membrane potential and adaptive threshold dynamics with two neuron cells, but omits internal induced-spike currents and sets reset voltage to resting potential.These changes reduce model generality, while retaining implementation of Eq. (9) biologically relevant spiking behaviors.
  • Results: 668.9 I&F neurons/mm2 were achieved, compared with 416.7 neurons/mm2 and 387.1 neurons/mm2 in two comparable 0.5 µm designs.A single neuron cell measures 41.7 µm × 35.84 µm and uses 62.3% of the area of the cited design in.
  • Results: The array showed mean output-to-input event ratios of 0.0208 ±1.22e-5 for 2040 M-N neurons and 0.0222 ±5.57e-5 for 4080 I&F neurons.The authors attribute the more controlled mismatch to shared synapse, comparator, and threshold-adaptation elements.
  • Results: 360 pJ per synaptic event was measured at 5.0 V, while operation at 1.0 V was validated more slowly by simulation and estimated at ~14.4 pJ per event.The 1.0 V value assumes dynamic energy scales with V2 while capacitance remains unchanged.
  • HiAER-IFAT: HiAER-IFAT extends AER routing with a multiscale tree for dynamically reconfigurable long-range connectivity, while external memory access remains the primary energy limitation.The limitation arises from DRAM table lookup for fully reconfigurable sparse synaptic connectivity.

3 DeepSouth

DeepSouth is a configurable FPGA cortex emulator that uses neurobiologically inspired hierarchical structure, event communication, and time multiplexing to simulate large spiking networks.

  • 3.1 Modular structure: DeepSouth uses a 100-neuron minicolumn as its fundamental computing unit and organizes connectivity hierarchically across neurons, minicolumns, and hypercolumns.Hypercolumns can contain up to 128 configurable minicolumns.
  • 3.2 Hardware implementation: A standalone Terasic DE5 implementation simulated 20 million to 2.6 billion LIF neurons in real time and 100 million to 12.8 billion at five times slower than real time.The architecture can also be implemented across multiple parallel FPGA boards.
  • 3.1 Modular structure: Time multiplexing lets one physical minicolumn simulate 200k time-multiplexed minicolumns, each updated every millisecond.The approach leverages FPGA speed and is limited mainly by available memory.
  • 3.1 Modular structure: Hierarchical connectivity stores neuron, minicolumn, and hypercolumn connection types instead of individual point-to-point connections, reducing communication cost by orders of magnitude.The design follows the structural organization observed in the cortex.
  • 3.1 Modular structure: Neural communication uses spike counts as events, allowing one minicolumn's events to reach up to 200k neurons through configurable hierarchical fan-out and delays.The 200k-neuron fan-out is realized across 16 hypercolumns, each containing up to 128 minicolumns of 100 neurons.
  • 3.2 Hardware implementation: The neural engine combines minicolumn, synapse, and axon arrays, while a Master coordinates memory, module progress, and external communication.The axon array applies programmable delays, and the synapse array weights events and assigns destination addresses.

4 BrainScaleS

BrainScaleS directly emulates neuron and synapse model equations with physical analog circuits and accelerates their temporal evolution, but its continuous-time implementation constrains single-chip scaling.

  • 4 BrainScaleS: BrainScaleS directly emulates model equations describing neuron and synapse dynamics, using electronic circuits whose electrical quantities represent model variables.The system was developed through collaboration among institutions in Heidelberg, Dresden, and Berlin.
  • 4 BrainScaleS: The first-generation BrainScaleS system implements the Adaptive Exponential Integrate-and-Fire model with linearly scaled parameters and accelerated model time.The membrane-voltage range between reset and firing threshold is approximately 500 mV.
  • 4 BrainScaleS: Because each BrainScaleS neuron and all its synapses use continuous-time analog circuitry, the system consumes substantial silicon area and requires multichip implementation beyond a few hundred neurons.The high acceleration factor also creates high communication-bandwidth requirements between ASICs, addressed through wafer-scale integration.
  • 4 BrainScaleS: A BrainScaleS wafer module integrates an uncut silicon wafer containing 384 neuromorphic chips.The wafer is shown beneath a copper heat sink in the module photograph.

4.1 HICANN ASIC

HICANN combines an analog AdEx neuron-and-synapse core with time-multiplexed digital event communication, configurable membrane circuits, and plastic synapses.

  • 4.1 HICANN ASIC: HICANN centers on a symmetrical analog network core in which synapse arrays enclose neuron blocks, each with associated analog parameter storage.The HICANN die also includes a surrounding communication network.
  • 4.1 HICANN ASIC: HICANN targets more than 10k presynaptic inputs per neuron and time-multiplexes event communication so each channel is shared by 64 synapses.Digital event signals correspond to biological action potentials.
  • 4.1 HICANN ASIC: HICANN communicates long-range neural events through packet switching with embedded digitized timing because accelerated wafer-to-host latencies cannot support real-time communication.A 100 ns communication latency corresponds to 1 ms of biological wall time at an acceleration factor of 10^4.
  • 4.1 HICANN ASIC: HICANN neurons are based on the AdEx model and use groups of interconnected membrane circuits to construct model neurons.Each membrane circuit has a column of 220 associated synapses, and spike generation can be enabled individually within a group.
  • 4.1 HICANN ASIC: Each membrane circuit receives two selectable synaptic input types, typically configured for excitatory and inhibitory inputs, using current-mode integration.An integrator restores the input voltage after each synapse sinks current for a nominal 4 ns interval.
  • 4.1 HICANN ASIC: Synapses select events by address matching, apply 4-bit DAC weights scaled by row-wise gmax, and support short-term plasticity and STDP.Short-term plasticity is implemented by modulating presynaptic enable-signal duration, while each synapse includes a correlation measurement circuit for STDP.

4.2 Communication infrastructure

HICANN's communication infrastructure provides dense, high-bandwidth L1 routing and wafer-scale interconnects to address the limits of conventional chip packaging and scaling.

  • 4.2 Communication infrastructure: The analog network core implements 220 synapse driver circuits, each receiving one L1 signal.L1 connections occupy routing space around and above parts of the analog network core.
  • 4.2 Communication infrastructure: 640 differential wires provide a total L1 bandwidth of 640 Gbits−1, with each wire pair supporting up to 2 Gbits−1.The wiring comprises 64 horizontal lines in the center and 128 lines on each side of the analog core, doubled for differential signaling.
  • 4.2 Communication infrastructure: Scaling beyond one HICANN requires chip-to-chip interconnect, while standalone 3D stacking can reliably stack only a few chips.Thus, 3D stacking alone is not sufficient for large network sizes.
  • 4.2 Communication infrastructure: Wafer-scale integration addresses both inter-reticle L1 continuity and wafer-to-PCB connections using post-processed metal and redistributed pads.The approach is compatible with silicon wafers manufactured in a standard CMOS process.
  • 4.2 Communication infrastructure: Post-processing uses 5×5µm2 minimum pad windows, 15×15µm2 copper areas, and 8.4µm interconnect pitch between adjacent reticles.The process has proven reliable down to a 6 µm pitch, although the implemented density was relaxed.

4.3 Wafer module

The BrainScaleS wafer module connects a wafer to its host and supports high-density electrical communication through precisely aligned contact structures. Its assembly uses Ethernet links, elastomeric connectors, alignment hardware, and resistance monitoring.

  • Host communication: The wafer module communicates with the host compute cluster through 48 1 Gbit/s Ethernet links.The links are provided by four I/O boards mounted on the communication subgroups.
  • Host communication: Future direct wafer-to-wafer networking is supported by additional connectors on the I/O boards.Twelve RJ45 connectors are visible on each described I/O board, while other connectors are reserved for future networking.
  • Wafer-to-PCB connection: The wafer connects to the main PCB through post-processed contact stripes with 1.2 mm width and 400 µm pitch.A matching mirror-image stripe pattern is placed on the main PCB.
  • Wafer-to-PCB connection: Elastomeric connectors with five conducting stripes per millimeter tolerate placement imprecision, requiring about 50 µm wafer-to-PCB alignment accuracy.An aluminum bracket, frame, and precision screws provide the alignment.
  • Assembly procedure: Assembly uses a slotted FR4 mask, optical inspection, and electrical-resistance monitoring to position connectors and achieve optimum contact pressure.The wafer is filled with material after correct placement, according to the assembly procedure.

4.4 Summary of the BrainScaleS system

BrainScaleS is designed as a wafer-scale, accelerated neuromorphic system combining continuous-time analog computation, programmable modeling, plasticity, and scalable integration. Compressing model time by several orders of magnitude enables learning and development simulations in seconds instead of hours.

  • System goals: BrainScaleS targets continuous-time analog neurons and synapses, programmable calibration and topology, per-synapse plasticity, wafer-scale scalability, and accelerated operation.These requirements are presented as simultaneous design goals for the system.
  • Acceleration: Several-orders-of-magnitude timescale compression enables processes such as learning and development to be modeled in seconds instead of hours.The authors connect this acceleration to parameter searches and statistical analysis across models.
  • Acceleration: The system is intended to make parameter searches and statistical analysis possible across neural models.This consequence follows the described accelerated analog neuromorphic operation.

5 Dynap-SEL: A multi-core spiking chip for models of cortical computation

Dynap-SEL is a mixed-signal, multi-core spiking processor combining analog neurons, programmable and plastic synapses, asynchronous event routing, and distributed memory. Its hierarchical routing and memory-optimized architecture support scalable multi-chip cortical-model implementations while reducing memory-access costs.

  • Architecture: Dynap-SEL combines four analog neural cores with programmable synapses and a fifth plastic core containing on-chip learning circuits.The four non-plastic cores contain 16×16 neurons and 64 programmable synapses per neuron; the plastic core contains 1×64 neurons and plastic and programmable synapses.
  • Routing: Three router levels handle intra-core, inter-core, and inter-chip communication using source-address and destination-address routing.The routing memory uses both SRAM and TCAM elements.
  • Non-plastic cores: Each non-plastic core contains 256 analog neurons and 16k asynchronous TCAM-based synapses arranged in a 2D array.Each neuron has 64 programmable-weight synapses with source-address tags, and TCAM supports 2^11 potential input sources per synapse.
  • External communication: Asynchronous AER communication connects Dynap-SEL chips with sensors and computing devices for large-scale sensory-processing systems.Parallel AER routing enables direct communication across chips and external devices.
  • Memory architecture: Memory and computation are co-localized in distributed SRAM and TCAM structures, reducing power and bandwidth demands relative to distant memory blocks.The architecture is described as eliminating the von Neumann bottleneck by design, while most current silicon area is occupied by memory cells.
  • Scalability: Plastic cores from up to 4×4 chips can be merged, enabling configurations with up to 1k neurons or 128×1k plastic synapses.Asynchronous memory control also permits online routing-table reconfiguration for structural plasticity or evolutionary algorithms.

6 The 2DIFWTA chip: a 2D array of integrate-and-fire neurons for implementing cooperative-competitive networks

The 2DIFWTA chip implements cooperative-competitive spiking networks with a dense 2D array of integrate-and-fire neurons, recurrent excitation, inhibition, adaptation, and configurable refractory behavior. Although designed for winner-take-all networks, it also supports broader neural modeling and real-time agent applications.

  • Network model: Cooperative-competitive networks use recurrent excitation and inhibition so strongly responding neurons suppress others while similarly tuned neurons cooperate.The chip was designed to study these networks in both mean-rate and time domains.
  • Chip architecture: The 2DIFWTA chip contains a 32×64 array of 2048 integrate-and-fire neurons with AER and local excitatory synapses.Its local connectivity supports either a two-dimensional network or 32 one-dimensional winner-take-all networks.
  • Neuron architecture: Each neuron integrates external excitatory and inhibitory AER inputs with local excitation and self-excitation through dedicated circuit blocks.The neuron includes AER input, local input, soma, and local output blocks.
  • Neuron model: The leaky integrate-and-fire circuit adds spike-frequency adaptation, a tunable refractory period, and voltage-threshold modulation while targeting low power.Adaptation models calcium-dependent after-hyperpolarization potassium currents and can reduce address-event communication bandwidth.
  • Applications: The chip’s general-purpose neuron-pool architecture enabled studies spanning olfactory processing, function approximation, inference, auditory perception, and latency coding.It was also used to demonstrate a real-time neuromorphic agent performing a context-dependent visual task.

7 PARCA :Parallel Architecture with Resistive Crosspoint Array

PARCA uses resistive crosspoint arrays to perform weighted sums and synaptic updates in parallel, combining analog computation within the crossbar with digital inter-array communication. Experimental implementations remain relatively small, and device non-idealities require circuit- and architecture-level mitigation.

  • Architecture: PARCA integrates resistive synapses into a crossbar to parallelize matrix-vector multiplication during read operations and synaptic weight updates during write operations.Peripheral circuits support both modes.
  • Read operation: Input-dependent read voltages interact with crosspoint conductances to produce weighted-sum currents at the column outputs.The input vector’s non-zero binary bits determine which read voltages are applied.
  • Read operation: Column read circuits integrate analog output currents, apply nonlinear activation such as thresholding, and convert results into spikes or digital values.Analog computation is confined to the crossbar core, while communication between arrays remains digital.
  • Write operation: Write-mode peripheral circuits generate programming pulses whose duty cycles reflect column neuron values, supporting gradient-descent or spike-based learning.Weight changes are represented as conductance changes in the resistive devices.
  • Implementations: Experimental crossbar demonstrations have generally used small arrays, including a 64×64 RRAM 1T1R neurosynaptic core with peripheral CMOS neurons.The 64×64 design monolithically integrates binary RRAM between M4 and M5 in a 130 nm CMOS process and time-multiplexes its column neurons.
  • Non-idealities: Circuit- and architecture-level strategies address off-state current, nonlinear updates, device variation, and IR drop in resistive arrays.Proposed measures include dummy columns, smart programming, multiple cells per weight bit, and relaxed wire widths.

8 Transistor-channel based programmable and configurable neural systems

Transistor-channel neuromorphic systems model biological channel dynamics with MOSFET-based circuits and combine dense synapses, configurable somas, learning, and event-based communication. These systems support applications including path planning and dendritic temporal classification while leveraging dense floating-gate synapses.

  • Transistor-channel modeling: Transistor-channel modeling exploits similarities between ion flow in biological channels and electron flow through MOSFET channels to implement dense soma circuits.The physical correspondence is approximate rather than identical.
  • Learning synapses: Floating-gate synapses store weights nonvolatily, generate biological EPSPs or PSPs, and support LTP, LTD, and STDP learning rules.The same device supports both long-term storage and synaptic response generation.
  • System architecture: A mesh or crossbar architecture interconnects configurable channel-neuron components, while AER blocks communicate action-potential events across the system.The IC combines synapses, soma circuits, and input/output spike infrastructure.
  • Density: Floating-gate approaches provide a synaptic-density advantage and are expected to scale with relatively similar density to EEPROM at a given process node.The comparison concerns working IC implementations and normalizes synapse area by the square of process-node size.
  • Path planning: 100% correct and optimal performance was reported across many randomized maze scenarios for path planning with realistic-neuron ICs.A wavefront propagated from the goal through the neuron grid to produce the solution sequence.
  • Dendritic computation: Dendritic computation was reported as 1000s of times more efficient than most analog signal-processing algorithms and was used for word spotting resembling HMM classification.The architecture also incorporates dendritic computation and FPAA reconfigurability.

9 Other neural emulators

Other neural emulators include Neurogrid, TrueNorth, and SpiNNaker, which use different mixed-mode or digital architectures for large-scale simulation and event-based processing.

  • Neurogrid: Neurogrid is a mixed-mode multichip system for large-scale neural simulation and visualization, modeling somas, dendrites, synapses, and axonal arbors.It contains 16 neurocores totaling 1 M neurons in sub-threshold analog circuits.
  • TrueNorth: TrueNorth provides 1 million digital spiking neurons across 4096 cores per die, with a single die consuming 72 mW.A 16-chip NS16e board consumes 1 W at 1 KHz speed.
  • SpiNNaker: SpiNNaker is a scalable digital neural array that uses brain-inspired communication for large-network simulation and event-based processing.Each node combines ARM968 processor cores, local memories, and a packet router.

10 Discussion

The discussion compares neuromorphic processors as systems with distinct strengths rather than ranking one as universally superior. Suitability depends on objectives and constraints such as biological fidelity, scalability, power, and density.

  • Overview: The reviewed simulators illustrate the field’s evolution and increasingly close the gap between brain-like computational efficiency and engineered systems.The paper presents the systems as state-of-the-art neural simulators with promising specifications and applications.
  • System comparison: PARCA uses resistive crosspoint arrays for parallel weighted sums and synaptic weight updates, while other systems emphasize distinct architectural capabilities.The discussion places PARCA within a broader comparison of event-based neural processors.
  • Trade-offs: Processor choice depends on objectives and constraints: analog systems more closely resemble biological neurons, whereas digital designs can offer robustness but may require more area and power.TrueNorth’s low digital power is attributed to its 28 nm process technology.
  • Trade-offs: IFAT, PARCA, and transistor-channel systems are identified as suitable when low power is the primary constraint because their designs were specifically optimized for it.The comparison also notes that different systems provide different advantageous properties.
  • Conclusion: The systems’ complementary advantages and disadvantages collectively contribute toward more energy-efficient, dense neural simulators that more closely mimic biological counterparts.The paper explicitly avoids deeming one system inferior to another.
Loading 1805.08932v1…