Source-linked AI summary

Scalable Optical Learning Operator

Uğur Teğin, Mustafa Yıldırım, İlker Oğuz, Christophe Moser, Demetri Psaltis

arXiv:2012.12404v2physics.opticscs.LGeess.IV

TL;DR

Existing optical learning systems need to combine linear optical computation with nonlinear elements and interfaces. This paper demonstrates a multimode-fiber framework using linear and nonlinear spatiotemporal mode interactions, achieving accuracy comparable to digital implementations across several learning tasks.

  • Problem

    Optical neural computation requires combining linear optical processing with nonlinear elements and input-output interfaces.

  • Method

    The framework uses simultaneous linear and nonlinear spatiotemporal interactions among spatial modes in multimode fibers as a learning computation engine.

  • Results

    Across image classification, speech recognition, and face-age prediction tasks, the system performs comparably to digital counterparts, including 93% diagnostic accuracy and 94.5% audio-digit accuracy.

  • Takeaways & Limitations

    Multimode-fiber spatiotemporal processing can support several learning tasks with accuracy comparable to digital implementations and better energy efficiency.

  • Takeaways & Limitations

    The authors replaced the initial CT-scan dataset because repetitive images and patient-level clustering could affect classification, using a 3000-image X-ray dataset instead.

Abstract

from arXiv · show

Today's heavy machine learning tasks are fueled by large datasets. Computing is performed with power hungry processors whose performance is ultimately limited by the data transfer to and from memory. Optics is one of the powerful means of communicating and processing information and there is intense current interest in optical information processing for realizing high-speed computations. Here we present and experimentally demonstrate an optical computing framework based on spatiotemporal effects in multimode fibers for a range of learning tasks from classifying COVID-19 X-ray lung images and speech recognition to predicting age from face images. The presented framework overcomes the energy scaling problem of existing systems without compromising speed. We leveraged simultaneous, linear, and nonlinear interaction of spatial modes as a computation engine. We numerically and experimentally showed the ability of the method to execute several different tasks with accuracy comparable to a digital implementation.

Introduction

SOLO combines linear and nonlinear optical processing in a shared multimode-fiber volume, using the fiber’s connectivity, interaction length, mode parallelism, and 2D interfaces. The framework supports versatile learning tasks with high speed and power efficiency, achieving accuracy comparable to digital computers.

  • Proposed approach: SOLO combines the linear and nonlinear parts of the optical system in a shared volume confined in a multimode fiber.The method is named SOLO, or Scalable Optical Learning Operator.
  • Proposed approach: The fiber combines optics’ 3D connectivity with long interaction length, lateral confinement, dense spatial modes, and a compact form factor.These properties enable relatively low-power optical nonlinearities while retaining high parallelism.
  • Processor architecture: The fixed, highly nonlinear multimode-fiber mapping is combined with a trainable single-layer digital decision network to create a reconfigurable processor.The optical output is recorded by a camera and used to train the decision layer on input-output pairs.
  • Demonstrated task: 93% accuracy was achieved for COVID-19 diagnosis from lung X-ray images using the fiber-produced representation and a trained single-layer classifier.A large database containing lung X-ray images with various diseases, including COVID-19, was used for training.
  • Performance: The nonlinear multimode-fiber mapping transforms input data into a nearly linearly separable output space at very high speed and power efficiency.The mapping differs from those used in the earlier support-vector-machine, reservoir-computing, random-mapping, and extreme-learning-machine approaches.
  • Scope and results: The framework covers regression, face-image age prediction, speech classification, and COVID-19 X-ray diagnosis, with accuracy comparable to digital computers.The paper also examines scaling to large data sizes and estimates power consumption per operation.

Learning a nonlinear function

The optical system learned a nonlinear Sinc-function mapping by transforming coded inputs through nonlinear fiber propagation before linear regression. Performance depended on pulse peak power, reaching minimum test error near 3.43 kW before deteriorating at higher powers.

  • Sinc-function regression: The task used inputs x between -π and π with labels y = Sin(πx)/(πx), a benchmark nonlinear regression relation.Each x was uniquely encoded as a 2D pattern for SLM recording.
  • Sinc-function regression: The system recorded nonlinearly propagated beam profiles for many inputs and applied linear regression to the resulting output data.The nonlinear propagation supplies the transformation needed before linear regression on the output representations.
  • Power dependence: At low peak power (~1.14 kW), performance was poor because the nearly linear mapping was not linearly separable.Transmission was nearly linear apart from the detector’s square-law response.
  • Power dependence: At around 3.43 kW, unseen test inputs achieved an RMSE of 0.0671, whereas further power escalation gradually deteriorated performance.Higher power drove propagation toward a Raman beam clean-up regime in which projected profiles became virtually unaffected by input data.

Abalone dataset

The framework was extended from interpolation to multivariable inference on the abalone dataset, predicting sea-snail age from eight physical parameters. It learned age from spatially distributed independent variables with an RMSE of 0.126.

  • Abalone dataset: The study tested the computing method on the abalone dataset for multivariable inference.This followed the recognition that interpolation alone is inadequate for complex inference problems.
  • Abalone dataset: Age prediction used eight physical parameters related to sea-snail age, including the number of rings.The parameters were recorded on the SLM as a 4x2 matrix with appropriate pixel scaling.
  • Abalone dataset: The recorded spatial distribution at the distal fiber facet was flattened into a 1D vector and fed to a decision layer for linear regression.This processing followed the spatial encoding on the SLM.
  • Abalone dataset: 0.126 RMSE quantified the framework’s accuracy in predicting abalone ages from spatially distributed independent variables.Figure 3 compared true ages with the corresponding predictions.

Face image dataset · Audio digit dataset

The face-image task predicted normalized age from face images, achieving an RMSE of 0.167 normalized years despite some impossible negative predictions. On audio data, SOLO classified spoken digits with 94.5% test accuracy and speakers with 95.2% after updating only the decision layer.

  • Face image dataset: SOLO estimated age from 9780 face images spanning ages 0–116, with ages normalized from 0 to 1.The dataset included people of different genders and ethnicities, and 1 represented 116 years.
  • Face image dataset: A single neuron served as the decision layer for age prediction using recorded fiber-output intensity profiles.The final regression layer produced some negative predictions, which are impossible ages.
  • Face image dataset: 0.167 normalized years was the achieved RMSE for age prediction.Predictions and true ages were shown for the first 1000 samples.
  • Audio digit dataset: Spoken-digit classification used English recordings from six people, converting one-dimensional audio into two-dimensional Mel spectrograms for SLM input.The spectrograms were encoded on high-peak-power pulses before output-intensity classification.
  • Audio digit dataset: 94.5% accuracy was obtained on test data for digit categorization using frequency-resolved beam-profile measurements.The decision layer classified the recorded fiber-output intensity images.
  • Audio digit dataset: SOLO measured the effect of nonlinear pulse propagation on categorization accuracy and reused the same audio dataset for speaker identification.Because the nonlinear transformation was task-independent, only the decision layer was updated.
  • Audio digit dataset: 95.2% accuracy was achieved on test data for speaker classification with frequency-resolved beam-profile measurements.Digital decision-layer loss and accuracy evolutions were presented alongside fiber-simulation results.

COVID-19 dataset

SOLO was tested on a difficult COVID-19 diagnosis task using 3000 X-ray samples. Classification in the decision layer achieved 83.2% accuracy on an unseen test set with frequency-resolved beam-profile measurements.

  • SOLO was evaluated for COVID-19 diagnosis using a dataset of 3000 X-ray samples.
  • X-ray samples were applied to pulses as phase modulation, and the corresponding fiber-output intensity patterns were recorded.
  • 83.2% accuracy was achieved on the unseen test set using frequency-resolved beam-profile measurement and decision-layer classification.

Physical model

The multimode fiber implements computation through sequential linear and nonlinear transformations of spatially encoded information. Modal dispersion, perturbation-induced mode mixing, and intensity-dependent coupling determine the propagation dynamics.

  • Linear propagation: In an ideal fiber at low power, modal and chromatic dispersion change mode phases at different rates without intermodal power exchange, producing a linear field transformation.This behavior is represented by the first term in Eq. 1.
  • Linear propagation: Fiber bending or impurities induce mode coupling that acts as linear mixing through matrix C.This contribution corresponds to the second term in Eq. 1.
  • Nonlinear propagation: At sufficiently high pulse peak power, nonlinear mode coupling performs a nonlinear operation on spatially encoded information throughout the fiber.This intensity-dependent contribution is represented by the third term in Eq. 1.
  • Nonlinear propagation: At each propagation step, fiber modes couple according to linear coefficients and a nonlinear coupling tensor η.The nonlinear operator multiplies each three-element combination of mode coefficients by the corresponding tensor entry.

Numerical studies

Numerical simulations reproduced learning with 70.8% COVID-19 X-ray accuracy but underperformed experiments because computational simplifications and fewer spatial modes limited nonlinear mapping. Simulations also showed that reducing pulse power or fiber length further decreased diagnosis accuracy.

  • Peak-power dependence: 68.8% and 67.2% diagnosis accuracy resulted when peak power decreased to half and quarter of the initial power, respectively.The numerically obtained baseline was 70.8% COVID-19 diagnosis accuracy.
  • Fiber-length dependence: 67.5% diagnosis accuracy resulted from reducing fiber length from 10 to 5 self-imaging periods, while a further reduction produced 64.5%.These simulations confirmed the importance of high-intensity light and sufficient propagation length for learning.
  • Computational cost: More than 2 years of GPU computation were required to simulate 3000 samples with exact experimental parameters.The reported computational burden motivated GPU-parallelized simulations and rescaling of propagation length and pulse peak power.

Experimental setup

The experiment used phase-shaped picosecond laser pulses launched into a 5 m graded-index multimode fiber, with camera-based or frequency-resolved output detection. Pulse conditions were optimized to preserve temporal unity and maximize spatial interactions, while categorization tasks showed performance increases.

  • Light source: A 10 ps, 125 kHz Yb fiber laser centered at 1033 nm provided the optical pulses.The laser output had a 10 nm spectral width.
  • Beam shaping: A phase-only SLM shaped the linearly polarized Gaussian beam before coupling it into the multimode fiber.The SLM had an 8um pixel pitch and 60 Hz speed, and used a grating phase pattern to expel unmodulated light.
  • Multimode-fiber platform: The optical engine comprised 5m of commercial GRIN 50/125 multimode fiber with NA 0.2 and 120 modes per polarization.A 15mm lens imaged the phase-modulated light onto the fiber, covering the whole core area.
  • Output detection: Fiber outputs were magnified 12.5 times and recorded with a 5.2um-pixel camera or measured frequency-resolved using a 600line/mm grating.The camera used a 4f imaging system, while the grating provided an alternative to 4f imaging.
  • Categorization tasks: Significant performance increases were observed for categorization tasks using the audio digit and Covid-19 datasets.This result is reported in the experimental setup passage without specifying numerical values.
  • Operating conditions: Pulse power and width were optimized to conserve temporal unity and maximize spatial interactions.Fiber input and output power were monitored continuously, and neutral-density filters prevented camera saturation.

Numerical Simulations

Numerical simulations used a GPU-parallelized time-dependent beam propagation method to model nonlinear pulse propagation in graded-index multimode fiber. The simulated outputs were converted into normalized intensity images, flattened, and linearly fitted for learning tasks.

  • Numerical setup: GPU-parallelized time-dependent beam propagation simulations modeled nonlinear pulse propagation through the fiber over 10 self-imaging periods.The launched pulses were centered at 1030 nm with one ps duration.
  • Numerical setup: For 3000-sample datasets, the propagation length was rescaled from 5 m to ~5.5 mm and peak power increased to 10 MW.This rescaling was used to reduce computing time while generating significantly nonlinear spatiotemporal evolution.
  • Data processing and readout: Data were encoded as multiplied beam phase information, then propagated pulses were time-averaged into normalized intensity images, flattened after downsampling, and linearly fitted.The final fitting used standard Linear Regression.

Supplementary Material of Scalable Optical Learning Operator

The supplementary material details SOLO’s multimode nonlinear and linear coupling model, numerical implementation, input encoding, and validation across regression and classification tasks. Simulations and experiments show accurate function learning, age prediction, audio classification, and robust measurements.

  • Model and implementation: SOLO models each mode’s nonlinear evolution with intermodal and intramodal coupling, while linear coupling represents perturbations such as bending and impurities.The GRIN-50/125 nonlinear coupling tensor has size 1204, and computing all nonlinear terms required two and a half months on a server computer.
  • Sinc-function learning: 0.0039 root-mean-squared error (RMSE) was obtained on test data when learning the Sinc function from simulated nonlinear propagation.Scalar inputs were expanded into two-dimensional form with a random mask, and the distal-end intensity distribution was used for linear regression.
  • Audio classification: Approximately 68% accuracy was obtained for simulated audio-digit classification, and 61% accuracy was obtained for speaker classification on unseen test data.Audio was converted into spectrograms and processed as a two-dimensional image-analysis task using fiber output beam shapes.
  • Robustness and measurement: Around 82% accuracy over test data was reproduced after a one-week interval, validating the robustness of the measurements.A supplementary measurement reported RMSE 0.079, corresponding to a 12.63 signal to noise ratio (SNR).

Supplementary References · Authors’ Note:

The revised manuscript updated the Audio digit and Covid-19 experiments after identifying ordered-measurement drift, replaced the Covid-19 CT-scan dataset with an X-ray dataset, and added frequency-resolved data collection. These revisions addressed dataset and measurement concerns while reporting improved classification accuracy and learning curves with frequency-resolved measurements.

  • Authors’ Note:: The revised arXiv preprint updated experiments and results for the Audio digit and Covid-19 datasets.
  • Authors’ Note:: Initial Audio digit and Covid-19 CT-scan measurements were ordered, allowing experimental drift to increase SOLO scores.
  • Authors’ Note:: The experiments were repeated in shuffled order, and the Audio digit and Covid-19 results were updated.
  • Authors’ Note:: A secondary frequency-resolved data-collection method using a diffraction grating was added alongside the 4F imaging method.
  • Authors’ Note:: Frequency-resolved measurement provided higher accuracy and better learning curves for classification tasks in the experiments.
  • Authors’ Note:: The initial Covid-19 CT-scan dataset was replaced with a Covid-19 X-ray dataset for the classification task.
  • Authors’ Note:: The CT-scan dataset contained 2482 scans from 120 patients, resulting in repetitive images and intrinsic clustering among data points.
  • Authors’ Note:: The replacement X-ray dataset contained 3000 images, with each image associated with a different individual.
Loading 2012.12404v2…