Source-linked AI summary
Single chip photonic deep neural network with accelerated training
Saumil Bandyopadhyay, Alexander Sludds, Stefan Krastanov, Ryan Hamerly, Nicholas Harris, Darius Bunandar, Matthew Streshinsky, Michael Hochberg, Dirk Englund
TL;DR
Photonic DNNs seek to reduce the computation, latency, and energy limits of electronic training, but integrating optical linear and nonlinear processing on one chip remains difficult. This paper builds a fully integrated coherent optical neural network with in situ hardware training and reports matched classification accuracy with a digital model, alongside projected high throughput and low energy per operation.
Problem
Efficient training of all-photonic DNNs with on-chip optical nonlinearities remains an open challenge despite the potential for low-latency optical inference.
Method
The FICONN combines coherent matrix multiplication, receiverless programmable nonlinear units, and in situ parameter optimization using hardware-measured directional derivatives.
Results
9.8 pJ/OP and approximately 0.53 TOPS per inference were estimated for the FICONN, while vowel-classification test accuracy matched the digital system.
Takeaways & Limitations
The demonstrated architecture supports single-chip optical inference and training, with a path toward devices that learn in real time for sensing, autonomous driving, and telecommunications.
Takeaways & Limitations
The proof-of-concept system is dominated by thermal phase-shifter power, while larger-scale efficiency improvements remain projected through alternative phase-shifter and NOFU technologies.
Abstract
from arXiv · showhide
As deep neural networks (DNNs) revolutionize machine learning, energy consumption and throughput are emerging as fundamental limitations of CMOS electronics. This has motivated a search for new hardware architectures optimized for artificial intelligence, such as electronic systolic arrays, memristor crossbar arrays, and optical accelerators. Optical systems can perform linear matrix operations at exceptionally high rate and efficiency, motivating recent demonstrations of low latency linear algebra and optical energy consumption below a photon per multiply-accumulate operation. However, demonstrating systems that co-integrate both linear and nonlinear processing units in a single chip remains a central challenge. Here we introduce such a system in a scalable photonic integrated circuit (PIC), enabled by several key advances: (i) high-bandwidth and low-power programmable nonlinear optical function units (NOFUs); (ii) coherent matrix multiplication units (CMXUs); and (iii) in situ training with optical acceleration. We experimentally demonstrate this fully-integrated coherent optical neural network (FICONN) architecture for a 3-layer DNN comprising 12 NOFUs and three CMXUs operating in the telecom C-band. Using in situ training on a vowel classification task, the FICONN achieves 92.7% accuracy on a test set, which is identical to the accuracy obtained on a digital computer with the same number of weights. This work lends experimental evidence to theoretical proposals for in situ training, unlocking orders of magnitude improvements in the throughput of training data. Moreover, the FICONN opens the path to inference at nanosecond latency and femtojoule per operation energy efficiency.
ARCHITECTURE
The FICONN processes neural-network inference coherently on chip by cascading programmable matrix transformations and nonlinear activations without intermediate optical-to-electrical conversion.
- Input encoding: The transmitter encodes each six-element input vector into optical amplitudes and phases across six channels.Each channel represents one input element in the optical field.
- Linear transformations: CMXUs transform optical field vectors through programmable unitary matrix operations implemented by MZI meshes.The transformation is represented as b^(n) = U^(n)a^(n).
- Nonlinear transformations: NOFUs apply programmable nonlinear activation functions to each transformed optical field before transmission to the next layer.The activation produces the next-layer input, a^(n+1) = f(b^(n)).
- End-to-end inference: Inference remains entirely optical between layers, with no photodiode readout, amplification, or digitization.Only the final output is read by an integrated coherent receiver.
- Output readout: The integrated coherent receiver homodynes the final output with a common local oscillator and converts it into a normalized classification distribution.The predicted class is the label with the highest normalized output value.
EXPERIMENT
The fabricated FICONN integrates optical transmission, coherent matrix multiplication, nonlinear activation, and coherent reception on one chip, while correcting hardware nonidealities and demonstrating low-loss, programmable operation.
- System performance: A 10 dB end-to-end loss across 91 optical-component layers corresponds to less than 0.1 dB insertion loss per component.This enabled single-shot inference without optical re-amplification.
- Input and matrix hardware: The transmitter fans input light into six channels that independently encode optical amplitudes and phases.A typical channel provides more than 40 dB extinction and supports input-vector programming with more than 13 bits of precision.
- Input and matrix hardware: The CMXU implements arbitrary 6 × 6 unitary operations using a 15-device MZI mesh.The mesh uses the Clements configuration.
- Matrix accuracy: 0.987 ± 0.007 average matrix fidelity was achieved after correcting hardware errors, losses, and thermal crosstalk, versus 0.900 ± 0.031 with direct programming.The comparison used 500 random 6 × 6 unitary matrices.
- Nonlinear hardware: 30 fJ per nonlinear operation was measured for the NOFU over a carrier lifetime of approximately 1 ns.The receiverless design directly connects the photodiode to the modulator, avoiding an intermediate amplifier.
- Nonlinear hardware: The programmable NOFU can realize different nonlinear functions by tuning cavity detuning and the tapped optical-power fraction.Its direct photodiode-to-modulator connection also removes amplifier-associated latency and power consumption.
IN SITU TRAINING
The FICONN trains weights and nonlinear activation parameters directly on photonic hardware using optically accelerated derivative estimation. This enables efficient in situ training for a fully integrated photonic DNN.
- Context: Photonic in situ training addresses prior methods that trained only linear layers or required digital gradient evaluation.Genetic-algorithm approaches were described as difficult to scale and requiring many generations to converge.
- Approach: In situ training optimizes both weights and nonlinear function parameters by evaluating their derivatives directly on hardware.The approach is described as robust to noise, average gradient descent, and convergence to a local minimum.
- Approach: The optimization perturbs all model parameters simultaneously along a random direction instead of perturbing one weight at a time.The update uses the directional derivative and learning rate η.
- Efficiency: Two training-set batches per iteration estimate the loss and directional derivative, reducing hardware evaluations compared with forward differences.The method also obtains true cost and derivative estimates despite component and calibration errors.
- Results: Over 96% training accuracy and over 92% test accuracy were achieved, with training using 16-bit weights.The system quickly exceeded 80% accuracy before asymptoting to 96% training accuracy.
DISCUSSION
The FICONN combines low-latency coherent optical computation with in situ training, while its current proof-of-concept performance is constrained by thermal phase-shifter power and implementation assumptions. The discussion outlines scaling paths for throughput, energy, and network types.
- Performance: 435 ps is the estimated FICONN inference latency, dominated by optical propagation through the PIC subsystems.The latency includes three CMXUs, two NOFUs, transmitter and receiver paths, and a U-turn.
- Performance: 9.8 pJ/OP is the estimated upper bound for on-chip energy consumption, with approximately 0.53 TOPS per inference.These estimates are reported in the supplementary information.
- Performance: Table I separates on-chip energy EOP from estimated total power Etotal,est and distinguishes predicted optimized-layout metrics from reported fabricated-layout latency.Predictions assume large batches and 50 GHz resonant modulators.
- Constraints: Thermal phase shifters dominate power consumption, requiring approximately 25 mW for a π phase shift.Alternative phase-shifter technologies are projected for varying neuron and layer counts.
- Implications: In situ training could accelerate energy-intensive model optimization and may provide noise-related regularization against overfitting.In the reported task, the FICONN and digital system had similar test performance, while the digital system overfit the training set.
- Scaling: Foundry fabrication could support larger systems, while spectral multiplexing and improved NOFUs could increase throughput and reduce energy.The discussion projects NOFU energy below 1 fJ/NLOP with alternative device technologies.
- Scaling: Recirculating waveguide meshes could extend the feedforward architecture to temporal or frequency data and feedback-based neural networks.The proposed extension would train phase-shifter settings in situ.
CONCLUSION
The work demonstrates a single-chip coherent optical DNN that performs inference and in situ training. Foundry fabrication and receiverless nonlinear processing support scaling, while direct hardware derivative estimation enables real-time learning applications.
- Conclusion: A single-chip FICONN performs both inference and in situ training using receiverless photodetection-driven nonlinear activation functions.This removes optical-to-electrical conversion between DNN layers and preserves phase information for optical data.
- Scaling: Scaling to hundreds of modes is projected to reduce energy consumption to approximately 10 fJ/OP while maintaining latencies orders of magnitude below electronics.The fabrication relies entirely on commercial foundry photolithography.
- Implications: Direct derivative estimation on hardware provides a generalizable in situ training approach for other photonic DNN architectures.The authors connect optically accelerated forward passes with real-time learning for sensing, autonomous driving, and telecommunications.
METHODS
The methods describe fabrication, characterization, calibration, and in situ training of the photonic integrated circuit, including software-based correction of hardware errors.
- System implementation: The PIC was fabricated in a commercial silicon photonics process, thermally stabilized, and electrically controlled through programmable current sources and custom drivers.The chip used fiber coupling, heatsinking, Peltier stabilization, 236 wirebonds, and buffered input transmission.
- System characterization: The transmitter, matrix meshes, nonlinear units, and receiver were characterized by sweeping phase-shifter currents and measuring optical transmissions or photocurrents.The first two meshes used NOFU photodiodes for output calibration, while the final mesh used the receiver.
- Hardware correction: 0.969 ± 0.023 average fidelity quantified how accurately the software model predicted hardware outputs.The model incorporated beamsplitter errors, waveguide losses, and thermal crosstalk, then supplied corrected hardware settings without real-time hardware optimization.
- System characterization: 0.22 ± 0.05 dB insertion loss per MZI corresponded to 1.32 ± 0.30 dB loss per CMXU.Losses were inferred by comparing transmission through paths containing different numbers of devices.
- In situ training: In situ training used six normalized vowel-formant features, 540 training samples, and 294 test samples with random parameter perturbations and directional-derivative updates.The experiments used δ = 0.05 and η = 0.002; a software feedback loop stabilized coupled optical power.
Supplementary Information: Single chip photonic deep neural network with accelerated
The supplementary information details MZI characterization and phase-shifter calibration procedures for constructing programmable coherent matrix multiplication units.
- MZI characterization: An MZI implements a programmable 2 × 2 unitary operation whose transmission is characterized at its bar and cross ports.The characterization sweeps heater current, measures voltage and transmission, and fits the device response.
- MZI characterization: The current-to-phase mapping is modeled with a fourth-order polynomial, θ1(I) = p4I 4 + p3I 3 + p2I 2 + p1I + p0.The fourth-order voltage-current fit accounts for non-Ohmic heater behavior at high currents.
- Internal phase-shifter calibration: CMXU calibration begins by optimizing main-diagonal phase shifters to maximize transmission from input 1 to output 6, deterministically setting them to the cross state.Remaining devices are calibrated by routing light through successive diagonals, using already calibrated MZIs to access subdiagonals.
- External phase-shifter calibration: External phase shifters are calibrated with meta-MZIs formed from neighboring MZIs programmed as 50-50 beamsplitters.Sweeping one external phase shifter and fitting output transmission determines relative static phases for the external heaters.
II. NONLINEAR OPTICAL FUNCTION UNIT
The NOFU uses carrier injection in a microring resonator to provide programmable nonlinear optical responses with picosecond-scale cavity dynamics.
- Resonator response: Increasing incident optical power injects carriers that increase round-trip loss, driving the resonator from overcoupling through critical coupling to undercoupling.The response therefore changes with injected photocurrent and incident power.
- Resonator response: Q ≈ 8300 corresponds to a 6.6 ps photon lifetime, so the cavity response does not limit device speed.The resonance was measured with no injected current.
- Power and modulation: About 75 µW detunes the NOFU by one linewidth under an assumed photodiode responsivity of approximately 1 A/W.At the experimental 0.8 V bias, operating power consumption is approximately 60 µW.
III. CORRECTING HARDWARE ERRORS
The FICONN corrects matrix-programming errors arising from component imperfections and thermal crosstalk in its thermally controlled phase shifters.
- Hardware error sources: Static component errors, transmission losses, and thermal crosstalk can reduce the accuracy of matrices programmed into the CMXU.The section motivates hardware error-correction procedures for improving matrix accuracy.
Transmitter correction
The transmitter’s thermal crosstalk is calibrated by measuring phase shifts caused by aggressor channels and correcting desired phase programs with the inferred crosstalk matrix. This correction improves phase-setting accuracy and repeatability, while external phase-shifter crosstalk is neglected.
- Measurement: Thermal crosstalk coefficients Mij are extracted by varying an aggressor channel and fitting the resulting static-phase change on another channel.The measured 12 × 6 matrix maps aggressor-channel settings to induced phase shifts.
- Scope: Crosstalk on the transmitter’s external phase shifters was neglected because coherent detection was unavailable directly at the transmitter output.The limitation applies specifically to those external phase shifters.
- Benchmark: 500 randomized programming experiments benchmarked whether correction could reliably set channel 1 to θ1 = π/2 while other channels varied.Implemented phase was inferred from measured transmission.
- Result: 0.493 ± 0.015 changed to 0.501 ± 0.003 for channel 2 after correction, demonstrating improved accuracy and repeatability.The comparison is reported for the measured phase on channel 2.
CMXU correction
CMXU crosstalk is difficult to isolate because mesh programming also redirects light and interacts with beamsplitter errors and loss. A fitted digital twin models these imperfections, while stochastic perturbation training provides average gradient descent for hardware optimization; the chip’s propagation time is about 435 ps.
- CMXU correction: CMXU crosstalk is difficult to measure directly because mesh programming changes phases and light paths while interacting with beamsplitter imperfections and device loss.These effects are difficult to disentangle from pure thermal crosstalk.
- Digital twin: A software digital twin fits beamsplitter errors, waveguide losses, and thermal crosstalk to measurements from the real device.The model treats these initially unknown imperfections as fit parameters.
- Digital twin: 0.969 ± 0.023 average fidelity was achieved when the software model predicted hardware outputs.The model was fit using 300 random unitary matrices and 100 random input vectors, then optimized with L-BFGS.
- Digital benchmark: 100% training accuracy versus 92.7% test accuracy shows that the digital benchmark overfits while matching the system’s test performance.The digital model and FICONN both achieved 92.7% on the test set.
- Stochastic optimization: Randomly perturbing all parameters yields gradient descent on average, avoiding the 2N model evaluations required per epoch by one-parameter finite differences.The expected update follows the gradient, with effective learning rate η = µ|π|/√N.
- Latency: 435 ps propagation time corresponds to the 29.8 mm transmitter-to-receiver waveguide path plus cavity lifetime.The latency estimate conservatively assumes ridge waveguides with ng ≈ 4.2.
Energy efficiency
FICONN inference combines coherent linear matrix operations and nonlinear activation functions on photonic hardware. For the demonstrated three-layer, six-mode system, the operation count and phase-shifter energy contribution are explicitly estimated.
- Operation count: 2MN^2 + 2(M −1)N operations per inference combine linear matrix operations with nonlinear activation functions.For large N, the linear term dominates and the count is approximated as 2MN^2.
- Energy estimate: Total energy per operation is estimated from latency and total photonic, driver, and readout power divided by the inference operation count.The power model includes phase shifters, nonlinear units, transmitters, and receivers.
- Operation count: 240 operations per inference result when M = 3 and N = 6.This count includes both linear and nonlinear operations.
- Energy estimate: 9.8 pJ/OP is the estimated phase-shifter contribution for the demonstrated system.The estimate uses 144 phase shifters, 37.5 mW average power per phase shifter, and 435 ps latency.
- Electronics: Electronic driver and readout energy is estimated from reported DAC, TIA, and ADC values assuming a 1 GHz clock rate.The clock assumption reflects the approximately 1 ns response time of the injection-mode NOFU.
Scaling
FICONN scaling favors larger mode counts, batching, and receiverless operation because electronics and nonlinear-unit costs are amortized across more optical modes. Throughput is ultimately constrained by input bandwidth, while higher clock rates trade latency against electronics power.
- Batching: Large input batches reduce effective latency because vectors can enter the PIC faster than the end-to-end propagation delay.For W vectors, latency is τlatency + (W −1)/fBW.
- Bandwidth: Input transmission rate, rather than end-to-end propagation time, determines ultimate speed and energy efficiency.The demonstrated carrier-injection NOFU has an approximately 1 ns response time; depletion-mode operation could improve bandwidth.
- NOFU scaling: A depletion-mode NOFU is estimated at 9 fJ/NLOP using 200 fF capacitance and a 0.3 V drive voltage.The estimate is based on CV^2.
- Scaling: 11.7 pJ/OP is estimated for the current system, while an optimized version reaches 513 fJ/OP.The optimized estimate assumes low-power phase shifters and high-speed electronics.
- Scaling: N = 34 modes reaches below 100 fJ/OP and N > 380 reaches below 10 fJ/OP for a receiverless three-layer system.With M = 10 layers, N = 10 modes is projected to outperform digital electronics in energy efficiency.
- Clock-rate trade-off: 50 GHz operation is assumed for low latency, but slower clocks may improve energy efficiency because electronics power scales nonlinearly with bandwidth.The optimal clock rate depends on whether latency or energy efficiency is prioritized.
- Receiverless scaling: Intermediate readout requires nearly twice as many modes as receiverless operation to achieve below 100 fJ/OP in a three-layer system.At 10 fJ/OP with M = 10, receiverless operation requires N = 114 versus approximately N = 575 with intermediate readout.