Source-linked AI summary
Delocalized Photonic Deep Learning on the Internet's Edge
Alexander Sludds, Saumil Bandyopadhyay, Zaijun Chen, Zhizhen Zhong, Jared Cochrane, Liane Bernstein, Darius Bunandar, P. Ben Dixon, Scott A. Hamilton, Matthew Streshinsky, Ari Novack, Tom Baehr-Jones, Michael Hochberg, Manya Ghobadi, Ryan Hamerly, Dirk Englund
TL;DR
Photon-starved optical neural networks face receiver-sensitivity and energy challenges that constrain edge deployment. Netcast decentralizes neural-network computation across photonic edge devices, achieving 98.8% accurate image classification with <1 photon per MAC and 3 THz bandwidth over deployed fiber.
Problem
Finite receiver signal-to-noise ratio limits optoelectronic neural-network operation in photon-starved environments.
Method
Netcast decentralizes neural-network computation using wavelength-parallel optical modulation and time-integrating receivers to reduce client energy consumption.
Results
98.8% accurate image classification was demonstrated with <1 photon per MAC and 3 THz of bandwidth over deployed fiber.
Takeaways & Limitations
The demonstrated architecture supports photonic edge computing over deployed fiber for applications including sensors, drones, and cellular networks.
Takeaways & Limitations
For multiple clients connected by fibers of different lengths, wavelength-dependent dispersion compensation cannot be used, limiting scalability across such links.
Abstract
from arXiv · showhide
Advances in deep neural networks (DNNs) are transforming science and technology. However, the increasing computational demands of the most powerful DNNs limit deployment on low-power devices, such as smartphones and sensors -- and this trend is accelerated by the simultaneous move towards Internet-of-Things (IoT) devices. Numerous efforts are underway to lower power consumption, but a fundamental bottleneck remains due to energy consumption in matrix algebra, even for analog approaches including neuromorphic, analog memory and photonic meshes. Here we introduce and demonstrate a new approach that sharply reduces energy required for matrix algebra by doing away with weight memory access on edge devices, enabling orders of magnitude energy and latency reduction. At the core of our approach is a new concept that decentralizes the DNN for delocalized, optically accelerated matrix algebra on edge devices. Using a silicon photonic smart transceiver, we demonstrate experimentally that this scheme, termed Netcast, dramatically reduces energy consumption. We demonstrate operation in a photon-starved environment with 40 aJ/multiply of optical energy for 98.8% accurate image recognition and <1 photon/multiply using single photon detectors. Furthermore, we show realistic deployment of our system, classifying images with 3 THz of bandwidth over 86 km of deployed optical fiber in a Boston-area fiber network. Our approach enables computing on a new generation of edge devices with speeds comparable to modern digital electronics and power consumption that is orders of magnitude lower.
EXPERIMENT
The experiment demonstrates Netcast using a silicon-photonic smart transceiver that transmits neural-network weights over 86 km of deployed optical fiber. The system achieves 8-bit precision and performs high-accuracy MNIST image classification after calibration.
- Experimental setup: 48 Mach-Zehnder modulators support modulation up to 50 Gbps, providing 2.4 Tbps of total bandwidth.The smart transceiver supports wavelength-division multiplexing with 16 lasers transmitting simultaneously at approximately −10 dBm (100 µW) per wavelength.
- Experimental setup: 86 km of deployed optical fiber carries the transmitted weights between the smart transceiver and the client.The deployed fiber connects the smart transceiver to the client in the experimental system.
- Results: The calibrated system performs image classification on the MNIST handwritten-digit benchmark.Weight data are encoded to multiple modulators simultaneously, and the experiment reports high accuracy for the task.
- Results: 8 bits of precision are achieved, exceeding the approximately 5 bits required for neural-network computation.The system is calibrated before image classification on a benchmark handwritten-digit task trained on a digital computer.
- Experimental setup: 16 laser sources operate across 3 THz of bandwidth with >25 dB optical SNR.The transceiver output uses wavelength-division multiplexing to combine separate modulator wavelengths.
ENERGY EFFICIENCY
Netcast minimizes client-side power by amortizing modulation and readout across many MACs. With near-term fan-out values of N = M = 100, the client can reach approximately 10 fJ/MAC, three orders of magnitude below possible existing digital CMOS.
- Client-side energy architecture: A single MZM and DAC encode input data across N wavelengths, enabling N MACs for each voltage applied to the modulator.Netcast also ensures each client component performs many MACs for modulation or electrical readout.
- Client-side energy scaling: ≈10 fJ/MAC is achievable for client energy consumption when N = M = 100.These are assumed near-term spatial and time-domain fan-out values.
- Comparison with digital CMOS: Three orders of magnitude lower than possible in existing digital CMOS is the projected client energy consumption.The scaling is summarized in Table I, which amortizes device energy by spatial fan-out N or time-domain fan-out M.
RECEIVER SENSITIVITY
Receiver sensitivity is constrained by photoreceiver noise, especially in photon-starved links, motivating operation near the shot-noise limit at approximately 1 photon per MAC. Time integration, reduced capacitance, and SNSPDs enable substantially lower-noise operation, including accurate computation below 1 photon per MAC.
- Noise limits: ≈1 photon / MAC is the target operating point for minimizing receiver noise in photon-starved optical links.Fiber propagation or free-space diffraction loss can force clients into photon-starved environments, where the lowest possible noise floor is needed.
- Time integration: Time-integrating receivers accumulate signal over M MACs, relaxing the required SNR compared with conventional amplified photoreceivers performing each MAC separately.The resulting SNR is M times higher than that from a single MAC of an amplified receiver.
- Thermal noise: ≈1fF-scale receivers lower thermal readout noise to the single-photon-per-MAC level by minimizing integration capacitance.The thermal noise floor is fundamentally linked to photodetector size, readout electronics, and their integration proximity.
- Single-photon detection: < 1 photon per MAC supports high-accuracy digit classification with superconducting nanowire single-photon detectors operating at the shot-noise limit.The receiver can use fewer than one photon per MAC because readout performs a vector-vector product with M = 100 MACs, producing a multi-photon measured signal.
- Security: Less than one photon per weight can prevent eavesdroppers from learning individual deployed weights, restricting them to mean statistics.This enables accurate computation while both the eavesdropper and client remain blind to the received weight data.
DISCUSSION
The discussion positions Netcast as deployable within networking hardware and extensible to programmable switches, while identifying photonic integration opportunities and fiber-dispersion constraints for scaling across clients.
- Deployment: Netcast’s cloud-based smart transceiver could reside in network switches, servers, or edge nodes and support programmable-switch inference.Prior work demonstrated layer-by-layer inference with smart transceivers.
- Deployment: Network-switch storage could hold multiple models for querying, while reflection-mode communication would request models using only a few bits.The querying communication may be slow and lossy.
- Energy scaling: ≈10 aJ/MAC receiver electrical energy is projected from emerging low-power and high-speed photonic technologies.Receiverless detectors, photonic DACs, and photonic ADCs could reduce energy further through tight transistor–photonic integration.
- Energy scaling: Existing commercial analog accelerators still consume watts of power despite promising lower neural-network power than electronic counterparts.Custom edge ASICs also face energy and bandwidth constraints comparable to larger CMOS processors.
- Limitations: Dispersion compensation using wavelength-dependent delays works for one smart transceiver and client but not when one transceiver serves clients over different fiber lengths.The discussion refers to supplementary material for the effects of dispersion.
CONCLUSION
The work presents a scalable photonic edge-computing architecture combining photonics and electronics, achieving orders-of-magnitude improvement over digital electronics. Demonstrations include photon-efficient image classification over deployed fiber and potential applications across internet-connected edge devices.
- CONCLUSION: The architecture combines photonics and electronics for scalable edge computing with orders-of-magnitude improvement over existing digital electronics.Its demonstrated components include wavelength-division multiplexing and time-integrating receivers.
- CONCLUSION: <1 photon per MAC and 3 THz of bandwidth enabled computing over deployed fiber with 98.8% accurate image classification.These results demonstrate operation at very low photon counts while using deployed optical infrastructure.
- CONCLUSION: The approach could support high-speed computing on deployed sensors and drones, live video processing on cellular devices and networks, and image classification on spacecraft.The proposed applications extend photonic computing across the internet’s peripheral nervous system and remote edge environments.
METHODS · Silicon Smart Transceiver
The silicon smart transceiver integrates 48 silicon photonic modulators for parallel optical modulation, using thermo-optic and free-carrier phase shifters to control bias and transmitted intensity. It was fabricated in a 130 nm-capable SOI foundry process for C-band operation and includes photodetectors.
- Silicon Smart Transceiver: 48 silicon photonic MZMs each support modulation at 50 Gbps.The transceiver contains 48 modulators, each capable of 50 Gbps modulation.
- Silicon Smart Transceiver: Each MZM uses two thermo-optic phase shifters to control bias and two free-carrier plasma dispersion phase shifters to control transmitted optical intensity.The two phase-shifter types provide separate bias-point and optical-intensity control.
- Silicon Smart Transceiver: The chip occupies 422 mm2 and connects to a printed circuit board through 336 wirebonds.These are the reported chip footprint and testing interconnect counts.
- Silicon Smart Transceiver: A 64 channel polarization-maintaining interface is used to couple light in and out of the chip.The passage describes alignment to a 64-channel polarization-maintaining coupling interface.
- Silicon Smart Transceiver: The 48 channel silicon smart transceiver was fabricated through the OpSIS IME foundry process multi-project wafer run.The fabrication used the OpSIS IME foundry process.
- Silicon Smart Transceiver: The IME foundry uses 248 nm lithography and can produce 130 nm CMOS electronics.The passage identifies the foundry’s lithography wavelength and CMOS capability.
- Silicon Smart Transceiver: The transceiver is designed for parallel modulation in the optical C-band at 1550 nm.Its optical operation is specified for C-band parallel modulation.
- Silicon Smart Transceiver: The SOI process forms 220nm-thick silicon components above 2 µm of buried oxide and includes photodetectors.These are the reported SOI layer dimensions and integrated detector capability.
Optical Energy Efficiency Measurement
The measurement setup benchmarked several photoreceivers, including calibrated superconducting nanowire single-photon detectors whose integrated output voltages mapped to discrete photon-count bins.
- Optical Energy Efficiency Measurement: Superconducting nanowire single-photon detectors were calibrated by integrating output voltage over a fixed window, with distinct voltage bins mapped to photon counts.The photoreceivers and calibration methods supported data generation for Figs. 4 and 5.
COMPETING INTERESTS
Several authors hold leadership or senior technical positions at companies involved in silicon photonics and optical computing, and some authors have filed patents related to Netcast; the remaining authors declare no competing interests.
- M.S leads silicon photonics at Nokia, while M.H, A.N, and T.B.J hold executive or technical roles at Luminous Computing, and D.B is chief scientist at Lightmatter.
- R.H and D.E filed a patent related to Netcast, while M.G, Z.Z, L.B, A.S, R.H, and D.E filed a related provisional patent.
- Other authors declare no competing interests.
CORRESPONDENCE · Supplementary Material: Delocalized Photonic Deep Learning on the Internet’s Edge
The correspondence identifies the authors and their institutional affiliations and introduces supplementary material on delocalized photonic deep learning. The material includes a section on mapping matrix-vector multiplication onto hardware.
- CORRESPONDENCE: The paper is authored by Alexander Sludds, Saumil Bandyopadhyay, Zaijun Chen, Zhizhen Zhong, and Jared Cochrane, among others.The author list continues across the correspondence pages.
- CORRESPONDENCE: The authors include Liane Bernstein, Darius Bunandar, P. Ben Dixon, Scott A. Hamilton, Matthew Streshinsky, and Ari Novack.These names appear in the continued author list.
- CORRESPONDENCE: The author list also includes Tom Baehr-Jones, Michael Hochberg, Manya Ghobadi, Ryan Hamerly, and Dirk Englund.These names complete the displayed correspondence author list.
- Supplementary Material: Delocalized Photonic Deep Learning on the Internet’s Edge: The affiliations include MIT’s Research Laboratory of Electronics and Computer Science and Artificial Intelligence Laboratory.Both MIT units are listed with Cambridge, Massachusetts addresses.
- Supplementary Material: Delocalized Photonic Deep Learning on the Internet’s Edge: The affiliations also include MIT Lincoln Laboratory, Nokia Corporation, and NTT Research Inc., PHI Laboratories.The listed locations are Lexington, New York, and Sunnyvale, respectively.
- Supplementary Material: Delocalized Photonic Deep Learning on the Internet’s Edge: The supplementary material contains a section titled “Mapping Matrix-Vector Multiply Onto Hardware” on page 19.This heading appears as section XV.
I. EXPERIMENTAL SETUP … IX. EFFECT OF SYSTEM LOSSES
The supplementary methods describe the photonic transceiver, client receivers, detector-noise modeling, calibration, bandwidth characterization, and system-loss analysis. Together, these measurements establish the implementation details and operating limits underlying the demonstrations.
- III. SINGLE PHOTON DETECTION EXPERIMENTAL SETUP: Single-photon experiments use superconducting nanowire single photon detectors, with comparator, buffering, attenuation, integration, and digitization electronics converting detector pulses into readouts.The setup uses a cryogenic WSi SNSPD and a custom voltage integrator based on a 10 pF internal capacitance.
- IV. CALCULATION OF RECEIVER NOISE; A. Thorlabs PDA10CS; B. Koheron PD100-DC; C. Thorlabs APD430C: Receiver-noise analysis compares amplified photodetectors, a linear avalanche photodetector, a custom time-integrating receiver, and an SNSPD using received optical power divided by the dominant noise source.Measured reference values include 300 µV RMS noise for the PDA10CS, 286 µV integrated noise for the PD100-DC, and 3 mV voltage noise for the APD430C.
- D. Time Integrating Receiver; V. TIME INTEGRATING RECEIVERS: The time-integrating receiver measures accumulated photocurrent by converting integrator-voltage changes into charge and then optical energy per MAC.Its measured RMS readout noise is ≈220 µV, matching the manufacturer specification, while the integration window in the main text uses M = 100.
- E. Superconducting Nanowire Single Photon Detectors: SNSPD calibration maps integrator voltage to photon number, and the shot-noise-limited experiment uses one MAC per readout because time integration does not reduce optical energy per MAC.The simulation adds Poisson noise to each multiply while sweeping the mean photon number.
- VI. CALIBRATION OF TIME INTEGRATORS; VII. TIME INTEGRATING RECEIVER BANDWIDTH: All 16 time integrators agree with the manufactured 10 pF integration capacitance, while their fundamental computation bandwidth is set by photodiode absorption spectra at 10’s of THz.The measured output voltage increases linearly with the number of optical pulses in an integration window.
- VIII. EYE DIAGRAM MEASUREMENT; IX. EFFECT OF SYSTEM LOSSES: Eye-diagram measurements characterize silicon modulators up to 10 GHz, while system losses can reach 30 dB per wavelength; a 25 µW client signal supports 250 GHz per wavelength at 100 aJ/MAC.The loss estimate assumes 10 dBm starting laser power, 70 km of deployed fiber, and 0.14 dB/km fiber loss.
X. DEPLOYED FIBER LINK … B. Shot Noise
The supplementary sections document Netcast’s deployed-fiber operation, loss compensation, photodetection, calibration, noise considerations, and scalability limits. Together, they report 98.8% classification across 3 THz over deployed fiber while identifying detector, dispersion, and nonlinear constraints.
- X. DEPLOYED FIBER LINK: 98.8% accurate classification across 3 THz demonstrates proof-of-principle simultaneous client computation over the deployed fiber.Commercial deployment could fill the available bandwidth using laser banks or comb sources with modulators and filters.
- XI. COMPENSATING LOSSES USING AN ERBIUM DOPED FIBER AMPLIFIER: 10 aJ/MAC of ASE noise reaches the receiver for a 100GHz channel when an EDFA provides gain of 100 to compensate ≈20dB losses.EDFAs increase bandwidth potential at fixed detector sensitivity but add amplified spontaneous emission noise.
- XII. NETCAST USING COHERENT DETECTION: Coherent detection can amplify weak received fields with a strong local oscillator, while locked combs reduce complex digital signal processing.A coherent comb scheme mixes receiver-encoded inputs with weight-encoded frequencies.
- XIII. RELATIVE INTENSITY NOISE: >30dB receiver SNR is predicted for mass-produced tunable lasers, implying laser relative intensity noise is not significant for Netcast.The estimate assumes -140dBc/Hz RIN and a 100GHz receiver wavelength channel.
- XIV. MODEL USED FOR IMAGE CLASSIFICATION: The image-classification model has layer sizes [784 →100 →100 →10], and finite hardware maps larger matrix-vector products across ⌈B/D⌉ integration windows.Calibration maps floating-point values to optical modulator voltages and decodes receiver signals; multiplication achieves ≈8 bits of accuracy.
- XVII. SUPERCONDUCTING NANOWIRE SINGLE PHOTON DETECTOR CALIBRATION: 30kHz single-photon operation keeps dark-count probability below 3% per integration window, while detector saturation occurs near 100fW or ≈10^6 photons per second.A linear voltage-to-photon mapping supports Poisson-statistics analysis and photons-per-MAC measurements.
C. Comparing Thermal Noise and Shot Noise … XXVII. ALTERNATIVE ARCHITECTURES FOR NETCAST
The supplementary sections characterize Netcast’s noise limits, deployment options, encoding strategies, and alternative architectures. They show feasibility across RF, satellite, polarization-drift, coherent-detection, and frequency-integrating implementations while identifying practical tradeoffs and limitations.
- C. Comparing Thermal Noise and Shot Noise: 1.6 ∗106 photons per readout marks the shot-noise crossover for a 10pF receiver, yielding an SNR of ≈1000, far above application requirements.The crossover scales linearly with capacitance, so lower capacitance makes shot-noise-limited operation easier, but shot-noise limitation need not be optimal.
- D. Dark Current: >10MHz operation makes the time-integrated signal dominate the detector’s quoted 50pA dark current, which is not observed.The detector is FGA01FC, with a quoted dark current of 50pA.
- E. Flicker ( 1 f ) Noise: 30kHz is the flicker-noise-to-shot-noise crossover for 1uA current and Kf = 10−8, below the approximately 100kHz readout bandwidth.The flicker-noise corner is expected to be lower for low-power time-integrating receivers operating near I ≈1uA.
- XXII. NETCAST AT RADIO FREQUENCIES: 241 power SNR and 15 voltage SNR are obtained for a 1GHz RF deployment bandwidth and 1µW local oscillator, sufficient for the applications.With time integration, 1pW received RF power and a 1µW local oscillator could enable 100GHz of accurate computation if sufficient spectrum exists.
- XXIII. DEPLOYMENT OF NETCAST TO SPACECRAFT: 10^16 MAC/s per wavelength and 10^19MAC/s across 100 wavelengths are estimated for low-earth-orbit deployment, while Mars deployment enables ≈10^9 MAC/s.The low-earth-orbit estimate assumes ≈1mW received per wavelength and 10−19J/MAC; the Mars estimate assumes ≈10−12 W per wavelength and shot-noise-limited detectors.
- XXIV. ENCODING NEGATIVE NUMBERS: Encoding negative values by shifting floating-point zero can amplify calibration errors through neural-network sparsity, creating the “fat zero” problem.Alternative encodings use separate wavelengths, polarizations, or spatial modes for positive and negative values, with bandwidth or hardware tradeoffs.
- XXV. POLARIZATION DRIFT AND MODE DISPERSION: A passive, broadband polarization-splitting scheme is proposed to resist polarization drift and mode dispersion without N = 100 active phase-stabilization components.The scheme splits and rotates polarization components, applies the same attenuation, and sums the detected photocurrents; duplicating receivers would double hardware and energy.
- XXVI. COHERENT DETECTION WITHOUT PHASE STABILIZATION; XXVII. ALTERNATIVE ARCHITECTURES FOR NETCAST: Sum-squared IQ detection removes phase and frequency dependence, while FITS uses wavelength-separated weight columns and can require only one or two photodetectors.These alternatives reduce stabilization requirements or change integration from time to frequency, but resonant filters can vary by ≈1nm after fabrication.