Source-linked AI summary

All-optical Nonlinear Activation Function for Photonic Neural Networks

Mario Miscuglio, Armin Mehrabian, Zibo Hu, Shaimaa I. Azzam, Jonathan K. George, Alexander V. Kildishev, Matthew Pelton, Volker J. Sorger

arXiv:1810.01216v3physics.app-phcond-mat.dis-nnphysics.optics

TL;DR

The work addresses speed and power limits in conventional electronic and electro-optic neural-network architectures caused by RC parasitic effects. It develops all-optical nonlinear activation functions using induced transparency and reverse saturable absorption, achieving nonlinear modulation and strong MNIST classification performance while suggesting photon time-of-flight operation for photonic networks.

  • Problem

    Electronic approaches remain constrained by RC parasitic effects that limit interconnect speed and power, motivating alternatives for high-speed, energy-efficient processing.

  • Method

    The study develops nanophotonic nonlinear activation functions based on Fano-resonance-induced transparency in plasmon-exciton assemblies and reverse saturable absorption in C60 films, then evaluates them in a three-layer MNIST network.

  • Results

    Approximately 3 dB nonlinear modulation is obtained for the plasmon-exciton system and approximately 7 dB for the C60 film, while optical activations reach approximately 100% MNIST accuracy except QD-Middle, which reaches 96%.

  • Takeaways & Limitations

    The proposed architecture suggests photonic neural networks could avoid parasitic switching and operate with runtime determined by photon time-of-flight through the network.

Abstract

from arXiv · show

With the recent successes of neural networks (NN) to perform machine-learning tasks, photonic-based NN designs may enable high throughput and low power neuromorphic compute paradigms since they bypass the parasitic charging of capacitive wires. Thus, engineering data-information processors capable of executing NN algorithms with high efficiency is of major importance for applications ranging from pattern recognition to classification. Our hypothesis is therefore, that if the time-limiting electro-optic conversion of current photonic NN designs could be postponed until the very end of the network, then the execution time of the photonic algorithm is simple the delay of the time-of-flight of photons through the NN, which is on the order of picoseconds for integrated photonics. Exploring such all-optical NN, in this work we discuss two independent approaches of implementing the optical perceptrons nonlinear activation function based on nanophotonic structures exhibiting i) induced transparency and ii) reverse saturated absorption. Our results show that the all-optical nonlinearity provides about 3 and 7 dB extinction ratio for the two systems considered, respectively, and classification accuracies of an exemplary MNIST task of 97% and near 100% are found, which rivals that of software based trained NNs, yet with ignored noise in the network. Together with a developed concept for an all-optical perceptron, these findings point to the possibility of realizing pure photonic NNs with potentially unmatched throughput and even energy consumption for next generation information processing hardware.

1. Introduction

Photonic neural networks could accelerate neural computation, but conventional optical nonlinearities rely on electro-optic conversion or detection that limits speed, power efficiency, and cascadability. This work proposes light–matter-interaction devices as all-optical activation functions for photonic neural networks.

  • Neural-network neurons combine weighted addition, nonlinear activation for data discrimination, and transmission to multiple destination neurons.
  • Photonic neural networks have demonstrated potential computing-speed increases of 2-3 orders of magnitude.
  • Conventional electro-optic absorption modulators introduce speed and power-efficiency trade-offs despite their controllability.
  • Optical-to-electrical-to-optical conversion hampers network speed and cascadability through charge-carrier movement and detector noise.
  • The study implements nanoparticle-based light–matter-interaction devices as nonlinear activation functions and compares them with Tanh, Sigmoid, and ReLu functions on specific deep-learning tasks.The proposed systems use induced transparency in a plasmon-exciton structure and reverse saturable absorption in C60 films.

2. Discussion

The study develops two all-optical nonlinear activation mechanisms and integrates them into photonic waveguides and neural-network evaluations. Induced transparency and reverse saturable absorption produce power-dependent transmission suitable for optical neural activation, with MNIST performance comparable to software-based activations.

  • Induced transparency: A CdSe quantum dot between gold nanorods forms a coupled plasmon-exciton system whose destructive interference produces induced transparency.The system is modeled as two coupled oscillators representing the quantum-dot and plasmon dipoles.
  • Induced transparency: As input power increases, saturation of the quantum-dot transition removes the Fano-resonance dip and makes energy dissipation a nonlinear function of power.The resulting response is spectrally narrow and is represented through the assembly’s extinction and imaginary refractive-index changes.
  • Waveguide integration: Embedding the nanoparticle/quantum-dot assembly in a waveguide yields a larger modulation range in the middle configuration than on top, reaching approximately 1.5 dB versus 0.2 dB.The two configurations were evaluated through transmitted power as a function of absorption and input power.
  • Array engineering: Arrays of nanoparticle/quantum-dot assemblies increase nonlinear transmission modulation to more than 2 dB on top and almost 3 dB in the waveguide middle.The arrays use closely spaced assemblies distributed in either waveguide configuration.
  • Reverse saturable absorption: C60 dispersed in PVA provides reverse saturable absorption modeled with five energy levels, and its nonlinear transmission produces approximately 7 dB modulation.A pump-probe analysis calculates transmission versus input fluence for a 1 µm-thick film with 10 mM C60 concentration.
  • Neural-network evaluation: All optical nonlinear activation functions achieve accuracy comparable to software-based networks, with approximately 100% validation accuracy except QD-Middle, which plateaus at 99%.QD-Middle reaches a maximum training accuracy of 96%, while validation curves are evaluated over the complete validation set.
  • System implications: The proposed architecture could reduce runtime to photon time-of-flight because it avoids parasitic switching and uses short optical delays.The paper estimates processing time from the physical length of the photonic integrated-circuit neural network.

3. Conclusion

The study demonstrates nonlinear optical responses from MNPs/QD and C60 platforms and applies them as activation functions in photonic neural networks.

  • The MNPs/QD system produces nonlinear modulation through interference between plasmonic nanoparticle dipoles and the quantum-dot exciton transition.
  • 3 dB is achieved for the MNPs/QD waveguide platform, while the C60 film provides approximately 7 dB modulation.
  • The nonlinear optical responses serve as activation functions in fully connected TensorFlow neural networks evaluated on MNIST handwritten-digit classifiers.
  • The ONN accuracy can match commonly employed alternatives for up to 50 reconfigurations during training and validation.

4. Methods

The simulations use finite-difference time-domain modeling to study coupling between the MNPs/QD system and a photonic waveguide.

  • FDTD Solutions models the MNPs/QD system and waveguide coupling by solving Maxwell equations on a discrete spatiotemporal grid.
  • The MNPs/QD absorptance is swept to model induced transparency as a function of input power.
  • An adaptive mesh algorithm refines the computational grid in the quantum-dot domain.

5. List of Abbreviations

The abbreviations list defines terms for nonlinear activation, optical neural networks, photonic integration, simulation, and related device concepts.

  • AF means Activation Function of the perceptron, and RSA means Reverse Saturable Absorption.
  • ADE means Auxiliary Differential Equation(s), and FDTD means Finite Difference Time Domain.
  • AONN means All-optical Neural Network, while PNN means Photonic Neural Network.
  • CMOS means Complementary Metal Oxide Semiconductor, and FPGA means Field Programmable Gate Array.
  • PIC means Photonic Integrated Circuit, and LMI means Light matter Interaction.

8. Article thumbnail upload

The passage describes a preview of the thumbnail image display on the author submission page.

  • The author submission page displays a preview of the thumbnail image.
Loading 1810.01216v3…