Source-linked AI summary

Deep Learning with Coherent VCSEL Neural Networks

Zaijun Chen, Alexander Sludds, Ronald Davis, Ian Christen, Liane Bernstein, Tobias Heuser, Niels Heermeier, James A. Lott, Stephan Reitzenstein, Ryan Hamerly, Dirk Englund

arXiv:2207.05329v1cs.ETphysics.optics

TL;DR

The paper addresses whether an optical neural network can incorporate matrix algebra and nonlinear activation while improving energy, density and latency over electronic hardware. It experimentally implements a coherent VCSEL-ONN using optical fanout, homodyne multiplication and inline nonlinear activation. The system reaches 7 fJ/OP and 25 TeraOP/(s·mm2), while also supporting rapid weight updating for training.

  • Problem

    Growing DNN models challenge electronic hardware, while ONNs still face high energy use, low compute density and latency from digitally implemented or absent inline nonlinear activation.

  • Method

    The paper experimentally implements a coherent VCSEL-ONN using micron-scale VCSEL arrays, spatial fanout, homodyne weighted accumulation and inline nonlinear activation.

  • Results

    7 fJ/OP full-system energy efficiency and 25 TeraOP/(s·mm2) compute density are achieved, with 93.1% MNIST digit-classification accuracy.

  • Takeaways & Limitations

    Rapid real-time weight updating and time-domain neuron encoding extend the system beyond inference toward training and models with billions of parameters.

Abstract

from arXiv · show

Deep neural networks (DNNs) are reshaping the field of information processing. With their exponential growth challenging existing electronic hardware, optical neural networks (ONNs) are emerging to process DNN tasks in the optical domain with high clock rates, parallelism and low-loss data transmission. However, to explore the potential of ONNs, it is necessary to investigate the full-system performance incorporating the major DNN elements, including matrix algebra and nonlinear activation. Existing challenges to ONNs are high energy consumption due to low electro-optic (EO) conversion efficiency, low compute density due to large device footprint and channel crosstalk, and long latency due to the lack of inline nonlinearity. Here we experimentally demonstrate an ONN system that simultaneously overcomes all these challenges. We exploit neuron encoding with volume-manufactured micron-scale vertical-cavity surface-emitting laser (VCSEL) transmitter arrays that exhibit high EO conversion (<5 attojoule/symbol with $V_π$=4 mV), high operation bandwidth (up to 25 GS/s), and compact footprint (<0.01 mm$^2$ per device). Photoelectric multiplication allows low-energy matrix operations at the shot-noise quantum limit. Homodyne detection-based nonlinearity enables nonlinear activation with instantaneous response. The full-system energy efficiency and compute density reach 7 femtojoules per operation (fJ/OP) and 25 TeraOP/(mm$^2\cdot$ s), both representing a >100-fold improvement over state-of-the-art digital computers, with substantially several more orders of magnitude for future improvement. Beyond neural network inference, its feature of rapid weight updating is crucial for training deep learning models. Our technique opens an avenue to large-scale optoelectronic processors to accelerate machine learning tasks from data centers to decentralized edge devices.

INTRODUCTION

The paper presents a coherent VCSEL optical neural network designed to address energy, density, latency, and scalability bottlenecks in neural-network hardware. Its experimental system combines coherent fanout, homodyne computation and inline nonlinearity, achieving high accuracy alongside strong full-system efficiency and compute density.

  • Nonlinearity: Homodyne detection supplies the network’s inline nonlinear activation with instantaneous response, avoiding digitally or optoelectronically implemented activation latency.The compute unit uses a VCSEL-specific nonlinear weighting function arising from coherent optical interference.
  • Architecture: The architecture uses a coherent axon VCSEL, optical fanout and weight VCSELs to perform parallel matrix operations with reduced device scaling.Sharing one input laser across j channels makes device requirements scale as O(j), compared with O(i × j) in the cited alternatives.
  • Experimental validation: 98% compute accuracy, corresponding to approximately 6 bits of precision, is demonstrated for homodyne vector multiplication.The reported accuracy is mainly limited by phase instability and the frequency response of injection-locked VCSELs.
  • System contribution: 7 fJ/OP full-system energy efficiency and 25 TeraOP/(s·mm2) compute density are reported for the VCSEL-ONN.These figures include digital electronics and represent 100× improvement compared to digital hardware.
  • Experimental validation: 93.1% MNIST digit-classification accuracy is achieved, exceeding 98% of ground truth while using more than 100 coherent VCSEL transmitters.The demonstrated model has a 28×28 input, two hidden layers, and a 10-channel output layer.
  • Scalability and training: Rapid real-time weight updating supports neural-network training, while time-domain neuron encoding is described as scalable to models with billions of parameters.The paper contrasts this programmability with the static weights used by almost all cited ONNs.

Supplementary Materials: Deep Learning with Coherent VCSEL Neural Networks

The supplementary materials identify the authors and their institutional affiliations in the United States and Germany.

  • The paper lists Zaijun Chen, Alexander Sludds, Ronald Davis, Ian Christen, Liane Bernstein, Tobias Heuser, Niels Heermeier, James Lott, Stephan Reitzenstein, Ryan Hamerly, and Dirk Englund as authors.
  • The authors are affiliated with MIT's Research Laboratory of Electronics and NTT Research in the United States.
  • Additional affiliations include Technische Universität Berlin's Fakultät II Institut für Festkörperphysik in Berlin, Germany.
  • The manuscript is identified as arXiv:2207.05329v1, dated 12 Jul 2022.

I. HOMODYNE DETECTION

The VCSEL-ONN uses coherent homodyne detection to perform vector multiplication and supports linear or nonlinear operation through amplitude or phase encoding. Weight phases encode matrix elements, while input modulation selects the operation type.

  • Homodyne architecture: Homodyne detection multiplies input and weight vectors through the interference term of two coherent VCSEL fields.Injection locking provides phase coherence, and a beamsplitter overlaps the two beams before photodetection.
  • Weight encoding: sin[φW(t)] ∝ Wij encodes weight-matrix elements through phase modulation of the weight lasers.The input vector and weight matrix are mapped across time steps and VCSEL transmitters.
  • Linear operation: Amplitude encoding of the input, AX(t) = Xi, activates linear multiplication XiWij in the homodyne product.The input laser phase is set to φX = 0 for the simplified linear response.
  • Nonlinear operation: Phase encoding of the input, sin[φX(t)] ∝ Xi, enables nonlinear operation with a programmable response.The phase of the weight laser tunes the homodyne nonlinearity, whose response is modeled in Supplementary Figure 1.
  • Detection: Balanced detection can improve signal-to-noise ratio, while a single detector offers greater system simplicity.The inference signal is contained in the AC term after direct-current components are coupled out.

II. ENERGY CONSUMPTION

The energy analysis identifies optical power needed for a target homodyne SNR as a fundamental consumption limit. It models system noise and uses that model to discuss current and future energy bounds.

  • Energy limit: Optical power required to achieve a desired homodyne SNR sets a fundamental limit on optical energy consumption.The SNR also determines the compute precision of the system.
  • Precision constraint: The section connects optical energy consumption to the precision requirement of homodyne computation.Higher required SNR corresponds to greater compute precision and therefore constrains the operating energy.
  • Energy model: The analysis models signal and noise sources, validates the model against experimental results, and evaluates current and future energy bounds.The stated model is applied to the lower bound imposed by the current system and to future improvement.

1. Signal-to-noise analysis

The SNR analysis models detector, shot, and laser-intensity noise in homodyne detection and compares the model with experiment. The measured SNR closely matches theory, with shot noise as the main noise contribution.

  • SNR formulation: SNR is defined from the homodyne signal relative to its noise amplitude and depends on input and weight laser powers.The power ratio is represented by γ = PW/PX, with PX denoting the input-laser power.
  • Noise model: The homodyne noise model includes detector thermal noise, photon shot noise, and laser intensity noise.The model uses detector NEP, optical powers, relative intensity noise, detection balance, and effective bandwidth.
  • Experimental validation: 135 measured SNR agrees with the theoretical prediction of 140 under the experimental conditions.The neural-network implementation reads out data at T = 10 ns per time step.
  • Noise contribution: Shot noise is the main contribution to the measured homodyne noise.The experimental result supports the model’s identification of the dominant noise source.

2. Analysis of system performance

System-performance analysis examines how integration time and optical power determine SNR. Under the modeled conditions, the system reaches useful precision at low photon counts and projects lower energy per operation with faster VCSELs and larger integration sizes.

  • Future performance: Less than 1 photon per operation is projected to support approximately 6–7 bits of compute precision.This precision estimate accompanies the high-speed VCSEL and long-integration scenario.
  • Integration strategy: Integrating over i time steps extends acquisition time to T = i·tc and improves SNR through longer integration.An integrating receiver reads the accumulated value only after the i time steps.
  • Noise regime: At low power and room temperature, thermal noise dominates the SNR rather than the shot-noise-limited regime reported for cooled detectors.The plotted power range has negligible laser intensity-noise contribution compared with detector and shot noise.
  • Experimental conditions: 100 SNR is obtained with 200 photons per operation, corresponding to 40 aJ/OP and about 7 bits of precision.The reported precision is described as sufficient for most neural-network tasks.
  • Future performance: At R = 25 GS/s and i = 10^6, less than 1 photon per operation is projected to enable 0.2 aJ/OP.The projection uses high-speed VCSELs and larger model sizes that integrate over a million neurons.

3. system power consumption

The system power analysis accounts for optical, electronic, memory, DAC, ADC, TIA, and integration costs. Low-voltage VCSEL modulation and multiplexing reduce energy per operation, while conventional DAC and memory components remain important costs.

  • ADC, TIA, and integration costs are reduced by triggering once after i time steps and dividing energy per use by 2·i operations.For image inputs, i=28×28 time steps correspond to integrating over a 28×28 image.
  • 3.7 aJ/sample is consumed by VCSEL electro-optic modulation at 1 GS/s.The modulation power is 3.7 nW with Vπ=4 mV.
  • 0.8 aJ/OP is the projected optical power consumption at R=25 GS/s with i=106 and j=32×32.Under these conditions, the laser consumes 40 µW electrical power and emits 10 µW optical power, fanning out to 10 nW per homodyne detector.
  • 3 fJ/OP is reached after each CMOS DAC activation is fanned out to 9×9 copies.Conventional DACs operate at 0.5 pJ/use, corresponding to 0.25 pJ/OP before fanout.
  • 1.6 aJ/use is required for DAC electro-optic modulation with C=200 fF/bit and V=0.4 mV.The proposed current-drive DAC uses low-voltage modulation matched to the VCSEL Vπ.
  • 100 fJ/access is required to deliver signals from memory through electronic wires.This cost is split across spatial fanout operations.

III. COMPARISON OF STATE-OF-THE-ART COMPUTING HARDWARE

The comparison evaluates computing hardware using energy efficiency, throughput density, and end-to-end energy efficiency. Reported reference systems provide baselines for assessing full-system performance.

  • 346 fJ/OP corresponds to 2.9 TeraOP/J for the cited digital-computing reference.
  • 0.27 TeraOP/s over 9 mm2 corresponds to the cited reference system's throughput density.
  • 14 pJ/OP is the cited reference system's end-to-end energy efficiency, equivalent to 0.07 TeraOP/J.
  • Full-system performance includes optical energy, nonlinear activation, readout, amplification, ADC, DAC, and memory access.

IV. FANOUT OF WEIGHTS

The architecture fans out encoded weight beams to process multiple input vectors simultaneously. A 5×5 VCSEL array can be split into 9×9 copies, producing 81 parallel input vectors and convolution-equivalent processing.

  • Weight beams encoding [W0, W1, ..., Wj] can be fanned out to k copies for simultaneous multiplication with k input vectors.
  • 9×9 phase-mask copies are produced from each VCSEL beam for spatial fanout.
  • The fanned-out 5×5 VCSEL-array operation is mathematically equivalent to convolution of the two input images.
  • 81 input vectors can be processed simultaneously with 9×9 fanout.

V. VCSEL ARRAYS

The VCSEL-array experiments characterize the fabricated hardware, optical wiring, bandwidth, resonance, and injection-locking operation. Fabricated devices reach 2 GHz bandwidth, while a commercial VCSEL enables 25 GS/s data encoding.

  • Array implementation: 400 VCSELs are included in one fabricated sample consisting of 16 arrays of 5×5 devices.The arrays are wire-bonded to a high-speed PCB for individual electrical control.
  • Array implementation: 4 mV is the applied drive amplitude matched to the VCSEL Vπ.Individual voltage biasing tunes VCSEL wavelengths for injection locking.
  • Bandwidth and resonance: 2 GHz is the measured 3 dB bandwidth of the fabricated VCSEL, compared with 7 GHz for a commercial VCSEL.
  • Bandwidth and resonance: 0.011(1) nm is the measured on-chip resonance linewidth, corresponding to a cavity Q factor of 105.
  • Bandwidth and resonance: The measured VCSEL bandwidth is 67% of the photon-lifetime limit.The discrepancy may reflect different cavity dynamics below and above the lasing threshold.

VI. INJECTION LOCKING

The section characterizes injection locking in VCSEL arrays, including its power-dependent locking range, frequency response, and π phase-shift voltage. Homodyne detection measures these properties and reveals thermal and free-carrier response regimes.

  • Array coupling: Diffractive optical elements fan out the leader laser to match the VCSEL-array geometry.The setup uses a 3×3 array with 80 µm beam spacing and a 5×5 array in supplementary characterization.
  • Locking characterization: Injection locking is characterized interferometrically by overlapping the leader laser with VCSEL beams and monitoring their beatnote.The beatnote is measured with a high-speed photodetector and spectrum analyzer; it shifts to DC when the VCSEL enters the locking range.
  • Frequency response: The injection-locked VCSEL frequency response is measured with homodyne balanced detection by increasing the modulation frequency at constant amplitude.The homodyne amplitude decays with frequency, and the free-carrier response above 10 MHz is 10 times weaker than the thermal response below 1 MHz.
  • Locking range: The injection locking range is proportional to the square root of input laser power.The VCSEL output power is fixed at 100 µW in this characterization.
  • Phase modulation: The measured π phase-shift voltage is Vπ=4 mV.The homodyne signal increases linearly with driving voltage until the VCSEL leaves injection lock at a peak-to-peak voltage of about 4 mV.

VII. DATA MODULATION AND DEMODULATION

The section uses local-oscillator modulation and demodulation to reduce thermal phase drift during data encoding. The recovered homodyne product agrees with the intended compute signal, with negligible residual error.

  • Data modulation: Thermal phase drifts at 1 GS/s are decoupled by mixing the data with a 2 GHz local oscillator.The modulated driving signal averages at zero, supporting thermal balance at each time step.
  • Data modulation and demodulation: Input and weight data are encoded by mixing with a local oscillator at twice the data rate, then demodulated to retrieve the compute signal.The local oscillator is used for both modulation and demodulation.
  • Compute verification: Demodulation of the sine wave produces an accurate homodyne product.The simulation result of Eq. 13 is compared with fNL2 to confirm correct signal demodulation.
  • Compute verification: The fNL1−fNL2 residual is 0.7% over 10,000 data points.This comparison verifies compute fidelity, and the reported residual is described as negligible.
Loading 2207.05329v1…