Source-linked AI summary
Survey of Machine Learning Accelerators
Albert Reuther, Peter Michaleas, Michael Jones, Vijay Gadepally, Siddharth Samsi, Jeremy Kepner
TL;DR
Machine-learning accelerators are proliferating across applications, architectures, and technology types, creating a need to compare their capabilities. This paper updates a prior survey by collecting public performance and power data, plotting accelerator trade-offs, and examining precision and training-versus-inference patterns. It finds denser and more diverse offerings, with many newer accelerators exceeding 1 TeraOps/W and training generally requiring at least 100W.
Problem
Accelerator technologies and publicly reported capabilities are rapidly expanding across applications, architectures, and numerical precisions, motivating an updated comparative survey.
Method
The paper compiles publicly available accelerator performance and power data, plots peak performance against power, and categorizes designs by precision, form factor, application, and inference or training capability.
Results
The survey finds more accelerator entries and architectural diversity, many newer designs above 1 TeraOps/W, and almost all sub-100W points representing inference-only processors.
Takeaways & Limitations
Accelerator competition spans vector, dataflow, neuromorphic, flash-based analog, and photonic technologies, with 8-bit integer reporting common for edge inference and 16-bit floating-point reporting common for data-center training.
Abstract
from arXiv · showhide
New machine learning accelerators are being announced and released each month for a variety of applications from speech recognition, video object detection, assisted driving, and many data center applications. This paper updates the survey of of AI accelerators and processors from last year's IEEE-HPEC paper. This paper collects and summarizes the current accelerators that have been publicly announced with performance and power consumption numbers. The performance and power values are plotted on a scatter graph and a number of dimensions and observations from the trends on this plot are discussed and analyzed. For instance, there are interesting trends in the plot regarding power consumption, numerical precision, and inference versus training. This year, there are many more announced accelerators that are implemented with many more architectures and technologies from vector engines, dataflow engines, neuromorphic designs, flash-based analog memory processing, and photonic-based processing.
I. INTRODUCTION
AI systems integrate sensing, data conditioning, algorithms, computing, robust AI, human-machine teaming, and users into end-to-end capabilities. This survey focuses on accelerators supporting computationally intensive DNN/CNN workloads, distinguishing training from inference, numerical precision, and emerging architectures.
- End-to-end AI systems connect sensors, data conditioning, algorithms, modern computing, robust AI, human-machine teaming, and users.
- Specialized accelerators are increasingly important as conventional scaling trends have ended and AI workloads demand different balances of performance and functional flexibility.
- The updated survey adds more accelerators and considers neuromorphic, memory-based analog, and light-based technologies while excluding most fixed FPGA and smartphone offerings.
- The survey emphasizes DNNs and CNNs because their computationally intensive dense matrix operations can exploit data-reuse architectures.
- Training adjusts model weights using labeled data, whereas inference applies trained weights to input data to produce predictions.
- Higher precision is generally used for training, while lower, especially integer, precision can be effective for inference.
- Neuromorphic accelerators commonly use synthetic neurons, synapses, and digitally encoded spiking signals; Tianjic supports selecting spiking or non-spiking layers per DNN layer.
II. SURVEY OF PROCESSORS
The survey compiles publicly available accelerator performance and power data and visualizes peak capability against power across application categories and technology choices. The resulting comparison shows denser, more diverse accelerator offerings, a broader precision range, and a persistent power distinction between training and inference.
- The survey gathers publicly available performance and power information and plots peak giga-operations per second against peak power.
- Figure 2 encodes numerical precision by geometric shape, form factor by color, and inference-only versus training-capable designs by hollow versus solid markers.
- Many newer accelerators exceed 1 TeraOps/W, while almost all points below 100W represent inference-only processors or accelerators.
- The survey reports a wider variety of released peak-performance numerical precisions, including exploration of limited and mixed precision.
- Five application categories span very-low-power, embedded, autonomous, data-center chip/card, and data-center system accelerators.
- Reported performance values may originate from frames-per-second benchmarks converted using model operation and memory information.
A. Research Chips
The survey reviews research accelerators spanning dataflow, sparse, memory-centric, neuromorphic, analog, and very-low-power designs. These chips illustrate architectural approaches for efficient inference across diverse power and application targets.
- Comparison framework: The survey compares publicly reported research-chip performance and power values using the same peak-performance-versus-power scatter-plot framework.
- Architectural approaches: Research chips include dataflow, sparse, memory-centric, and neuromorphic architectures for efficient neural-network processing.NeuFlow uses matrix-multiplier and convolver tiles; EIE exploits sparsity and compressed models; TETRIS places weights and data near processing elements; TrueNorth uses digital spiking neural networks.
- Neuromorphic research: The TrueNorth chip draws up to 275 mW, while its system draws 44 W.
- Emerging technologies: Recent research and commercial examples extend inference acceleration to analog flash memory, processor-in-memory, and concurrent multimodal processing.Mythic uses flash-based variable resistors for matrix multiplication; Perceive targets concurrent video and audio networks; Syntiant performs int4 weight and int8 activation computation in processor memory.
C. Embedded Chips and Systems
Embedded accelerators combine CPUs, coprocessors, dataflow engines, and processor-in-memory techniques for high-performance inference under constrained power and form-factor requirements.
- Application scope: Embedded systems target high-performance inference for cameras, small UAVs, modest robots, and related devices.
- Embedded architectures: ARM Ethos N77 integrates four MAC Compute Engines, each capable of 1 TOP/s at 1.0 GHz.
- Embedded architectures: The BM1880 combines two ARM Cortex CPUs, a RISC-V controller, and a tensor processing unit for video surveillance.
- Embedded architectures: Gyrfalcon’s Lightspeeur 5801 uses a matrix engine and processor-in-memory techniques for model inference, with a server incorporating 128 chips.
D. Autonomous Systems
Autonomous-system accelerators target automotive, UAV, and robotic inference through heterogeneous CPUs, GPUs, NPUs, FPGA-based designs, and dataflow engines.
- Application scope: Autonomous-system entries target inference for automotive AI/ML, autonomous vehicles, UAVs, remotely piloted aircraft, and robots.
- Programmable accelerators: AImotive aiWare3 is a programmable FPGA-based accelerator aimed at autonomous driving.
- Architectural diversity: Other systems include agent-based computation, integrated CPU-AI accelerators, VLIW cores with AI coprocessors, and GPU-equipped platforms.The surveyed examples include AlphaIC RAP-E, Huawei Ascend 310, Kalray Coolidge, NVIDIA Jetson and Xavier, and Quadric accelerators.
- Reported measurements: Hailo-8’s plotted peak-performance value comes from ResNet-50 inference because no peak-performance number or derivable architectural details were published.
- Heterogeneous systems: Tesla’s FSD Computer uses two NPUs with 96x96 MAC arrays, alongside 12 ARM Cortex-A72 CPUs and a GPU.
E. Data Center Chips and Cards
The data-center chips and cards category groups CPUs, GPUs, programmable FPGA solutions, and dataflow accelerators for larger-scale machine-learning processing.
- Technology categories: Data-center chips and cards include CPUs, GPUs, programmable FPGA solutions, and dataflow accelerators.
1) CPU-based Processors: •
The surveyed CPU- and FPGA-related accelerators span conventional processors, reconfigurable hardware, and programmable inference cards, with varied precision and deployment targets.
- CPU-based processors: The second-generation Xeon processors are versatile inference engines, with peak values computed from dual AVX-512 units across two CPUs.The 8180 configuration is used for inference and the 8280 configuration for training.
- CPU-based processors: The Pezy SC2 combines 2,048 cores with eight threads per core for dense linear algebra suitable for AI training and inference.
- CPU-based processors: Tenstorrent Grayskull supports int8, fp16, and bfloat16, while its floating-point formats operate at one-quarter of int8 performance.
- CPU-based processors: The PFN-MN-3 card contains four chips with four dies each, 32 GB of RAM per chip, and fp16 matrix arithmetic across each die.
- FPGA-based accelerators: FPGA-based offerings emphasize programmable inference, including Arria CPU-FPGA processing, VectorPath int8 computation, Flex Logix multi-precision support, and Microsoft Brainwave reprogrammability.Training is described as more challenging for the Arria CPU-FPGA approach, while Cornami’s prototype has reported fp16 performance but its ASIC has not taped out.
3) GPU-based Accelerators:
The surveyed GPU, dataflow, and specialized accelerator entries cover both training and inference, with architectures ranging from general-purpose computation cards to statically routed processors and wafer-scale engines.
- GPU-based accelerators: V100, A100, MI8, and MI60 GPUs target both inference and training, whereas the T4 is geared primarily toward inference.
- Dataflow chips and cards: Alibaba’s Hanguang 800 posted the highest inference rate for a chip when announced, although only a ResNet-50 benchmark was reported.
- Dataflow chips and cards: Cambricon reported significant results for both int8 inference and float16 training, so both measurements appear on the chart.
- GPU-based accelerators: Google’s TPU1 supports inference only, while TPU2 and TPU3 support both training and inference.
- Dataflow chips and cards: Groq’s TSP contains over 400,000 multiply-accumulate units organized into Superlanes executing statically scheduled VLIW programs.Its processing units operate on 8-bit words and can be grouped for floating-point and multi-byte integer computation.
- Dataflow chips and cards: The Cerebras WSE places over 400,000 cores across 84 chips on one wafer and draws a maximum of 20kW.Its plotted performance is an estimate because clock speed, precision, and computational performance were not released.
F. Data Center Systems
The surveyed data center systems combine multiple accelerator cards or specialized processors into larger single-node platforms, with reported or estimated system-level power and performance.
- Data center systems: The section covers single-node data center systems rather than individual accelerator chips or cards.
- Data center systems: NVIDIA’s systems range from the four-V100 DGX-Station to DGX-1 with eight V100 GPUs and DGX-2 with sixteen V100 GPUs.The DGX-Station is a tower workstation, while DGX-1 and DGX-2 are rack-mounted servers.
- Data center systems: GraphCore’s DSS8440 IPU-Server contains eight C2 cards, with server power estimated from a typical dual-socket Intel server and eight PCI cards.
- Data center systems: The SolidRun Janux GS31 incorporates 128 Gyrfalcon Lightspeeur 5801 chips and uses processor-in-memory techniques for inference at a maximum 900W.
III. ANNOUNCED ACCELERATORS
The announced-accelerator landscape includes prospective conventional, neuromorphic, biological, and photonic designs, but many entries lack enough public information for performance or power assessment.
- Information boundaries: SambaNova, Blaize, and Esperanto had not released enough architectural, performance, or power information for quantitative assessment.
- Neuromorphic and photonic designs: Eta Computing shifted from its TENSAI spiking-neural-network demonstration toward a conventional inference chip after judging spiking chips not ready for commercial release.The commercial release status of TENSAI remained unclear.
- Neuromorphic and photonic designs: Announced neuromorphic efforts include BrainChip Akida, Anaflash, NeuronFlow, biological-neuron circuitry from Koniku, and Intel Loihi.Loihi had been scaled to 768 chips simulating 100 million spiking neurons, while Akida features 1,024 neurons per chip and less than one Watt of power.
- Neuromorphic and photonic designs: Several companies announced optical or silicon-photonic AI processing, while the survey deferred adding entries until performance and power numbers become available.
- Industry status: The survey also notes canceled accelerator programs and bankrupt companies, including Intel’s halted Nervana NNP-T development and curtailed NNP-I development.
IV. SUMMARY
The survey spans deep neural network accelerators from extremely low-power embedded systems to data-center platforms for inference and training. It also reports broader architectural diversity and differing precision practices across deployment settings.
- The survey covers accelerators ranging from extremely low-power embedded and autonomous applications to data-center inference and training.
- Many newly announced accelerators report peak performance and power consumption data.
- Edge and embedded accelerators increasingly report 8-bit integer performance, whereas data-center training accelerators often report 16-bit floating-point performance.
- The surveyed landscape includes neuromorphic, flash-based analog-memory, dataflow, and photonic-processing architectures.