Source-linked AI summary
Training of photonic neural networks through in situ backpropagation
Tyler W. Hughes, Momchil Minkov, Yu Shi, Shanhui Fan
TL;DR
Integrated photonic neural networks efficiently perform matrix operations, but lack an efficient training protocol that operates on the physical device. The paper introduces TRIM, which implements adjoint-variable backpropagation through in situ intensity measurements, and demonstrates training in a simulated photonic ANN. The method also applies to broader photonic sensitivity analysis and reconfigurable-optics optimization.
Problem
Integrated photonic neural networks lack an efficient in situ training protocol, while model-based and brute-force approaches are inefficient or limited by physical-model accuracy.
Method
TRIM physically implements the adjoint variable method to compute photonic ANN cost-function gradients from in situ intensity measurements without an external system model.
Results
91% training accuracy and 91% test accuracy were achieved after around 4000 iterations in a numerically simulated photonic ANN.
Takeaways & Limitations
The method supports efficient, scalable in-device training and may extend to sensitivity analysis and optimization of reconfigurable photonic systems.
Takeaways & Limitations
The method is exact for lossless, feed-forward, reciprocal systems; uniform loss requires gradient rescaling, while mode-dependent loss limits accurate time-reversed-field reconstruction.
Abstract
from arXiv · showhide
Recently, integrated optics has gained interest as a hardware platform for implementing machine learning algorithms. Of particular interest are artificial neural networks, since matrix-vector multi- plications, which are used heavily in artificial neural networks, can be done efficiently in photonic circuits. The training of an artificial neural network is a crucial step in its application. However, currently on the integrated photonics platform there is no efficient protocol for the training of these networks. In this work, we introduce a method that enables highly efficient, in situ training of a photonic neural network. We use adjoint variable methods to derive the photonic analogue of the backpropagation algorithm, which is the standard method for computing gradients of conventional neural networks. We further show how these gradients may be obtained exactly by performing intensity measurements within the device. As an application, we demonstrate the training of a numerically simulated photonic artificial neural network. Beyond the training of photonic machine learning implementations, our method may also be of broad interest to experimental sensitivity analysis of photonic systems and the optimization of reconfigurable optics platforms.
I. INTRODUCTION
Integrated photonics can accelerate neural-network matrix operations, but efficient in situ training remains unresolved. The paper proposes TRIM, which uses adjoint-variable methods and intensity measurements to compute gradients efficiently.
- Photonic circuits are attractive for neural networks because they efficiently perform the matrix-vector multiplications used extensively in these models.
- Efficient on-platform training is essential, yet integrated photonic networks have relied on computer models or inefficient brute-force gradient estimation.Model-based training depends on physical-model accuracy, while serially perturbing each parameter scales poorly.
- TRIM computes photonic ANN cost-function gradients using only in situ intensity measurements.The procedure physically implements the adjoint variable method.
- The proposed method scales in constant time with the number of parameters, enabling efficient backpropagation in hybrid optoelectronic networks.The derivation begins from Maxwell’s equations and may extend beyond the specific hardware implementation discussed.
- The paper derives the forward and backward propagation mathematics, explains the physical gradient procedure, validates it numerically, and demonstrates photonic-network training.
II. THE PHOTONIC NEURAL NETWORK
The photonic ANN alternates programmable linear optical operations with electronic nonlinear activations, and its gradients are propagated backward through the circuit. The training formulation connects these gradients to tunable phase-shifter parameters and physically measurable fields.
- Network operation: A feed-forward photonic ANN alternates linear matrix operations with element-wise activations and minimizes a cost function over training examples.Backpropagation computes the required gradients by applying the chain rule from the output layer toward the input layer.
- Network operation: Optical interference units use controllable Mach–Zehnder interferometers to implement unitary N × N operations, while nonlinear activations run in an electronic circuit.The optical state is measured before electronic activation and prepared for injection into the next ANN stage.
- Network operation: The OIU is modeled as a linear, lossless, forward-only device connecting N single-mode input ports to N single-mode output ports.Its transfer matrix is the off-diagonal block of the full scattering matrix, with reciprocity represented by transpose-related blocks.
- Forward and backward propagation: Forward propagation applies each layer’s matrix ˆW_l followed by its element-wise activation, beginning with input X0 and ending at layer L.
- Gradient computation: Training minimizes the cost with respect to linear operators ˆW_l by tuning integrated phase shifters, but existing OIU-tuning methods do not directly provide ANN gradients.
- Gradient computation: The proposed gradient method avoids an external system model and tunes parameters in parallel, unlike brute-force computation, which scales linearly with parameter count.
- Gradient computation: The derivation accounts for both the explicit dependence of ˆW_l on a phase-shifter permittivity and the implicit dependence of later-layer fields on that parameter.
- Forward and backward propagation: The backward pass recursively computes δ_l vectors from the output toward the input, physically propagating them and Γ_l through the OIUs.The cost function is demonstrated as a mean-squared function of the output relative to a complex-valued target vector.
III. GRADIENT COMPUTATION USING THE ADJOINT VARIABLE METHOD
The adjoint variable method expresses photonic ANN cost-function gradients through the overlap of original and adjoint electromagnetic fields, enabling parallel computation for all phase shifters.
- The OIU’s transfer matrix relates input and output modal amplitudes, while its port-source formulation connects those amplitudes to electromagnetic field solutions.
- Maxwell’s equations provide the electromagnetic formulation used to compute gradients with respect to the physically adjustable phase-shifter permittivities.The formulation uses the spatial relative permittivity, electric field, current source, and a symmetric operator arising from Lorentz reciprocity.
- The ANN gradient with respect to phase-shifter permittivities can be expressed using the original and adjoint electromagnetic field solutions.The original field is generated by the network input, while the adjoint field is associated with the cost-function derivative.
- The gradient is given by the overlap of the original and adjoint fields over the phase-shifter positions.
- The resulting formulation computes the loss gradient in parallel for all phase shifters when the original and adjoint fields are known.
IV. EXPERIMENTAL MEASUREMENT OF GRADIENT
TRIM obtains photonic ANN gradients through in situ intensity measurements by combining original and time-reversed adjoint fields inside the device.
- TRIM computes the gradient of a photonic ANN using intensity measurements taken within the optical device.The method generates an interference pattern whose cross term contains the desired gradient information.
- The adjoint field is sourced from the device’s output ports, and time reversal generates the conjugated adjoint field needed for interference.
- The time-reversed adjoint input amplitudes can be determined experimentally by sending δl into the output ports and conjugating the measured input-port result.
- The experimental procedure measures original-field intensities, sends adjoint amplitudes through the output ports, and records the corresponding phase-shifter intensities.
- The original and time-reversed adjoint fields are interfered, after which constant intensity terms are subtracted and the result is multiplied by k0^2 to recover the gradient.
V. NUMERICAL GRADIENT DEMONSTRATION
The TRIM method reconstructs photonic adjoint sensitivities from in situ intensity measurements, matching direct adjoint calculations with high precision despite simulated losses.
- V. NUMERICAL GRADIENT DEMONSTRATION: The original field and time-reversed adjoint field are interfered, with constant intensity terms subtracted to obtain gradient information.The reconstructed pattern uses the field from panel (b) and the time-reversed adjoint field from panel (d).
- V. NUMERICAL GRADIENT DEMONSTRATION: The simulated three-MZI system compares direct adjoint gradients with gradients reconstructed by TRIM.The setup uses a 3×3 linear operation and evaluates fields and permittivity sensitivities across the device.
- V. NUMERICAL GRADIENT DEMONSTRATION: Panels (e) and (f) match with high precision, demonstrating agreement between direct adjoint and TRIM gradient calculations.The comparison concerns the gradient of the objective function with respect to the permittivity at each spatial point.
- V. NUMERICAL GRADIENT DEMONSTRATION: The simulated system has 41% total power loss from scattering and approximately 5–10% mode-dependent loss, yet TRIM retains very good sensitivity fidelity.The losses arise from sharp bends, stair-casing, and differences associated with injection at different input ports.
VI. EXAMPLE OF ANN TRAINING
The paper numerically trains a two-layer photonic ANN to implement XOR using in-situ backpropagation and adjoint gradients. The small demonstration learns the nonlinear mapping from four training examples.
- VI. EXAMPLE OF ANN TRAINING: The demonstration trains a photonic ANN to implement XOR, mapping [0 0]T and [1 1]T to 0, and [0 1]T and [1 0]T to 1.XOR is used as a simple nonlinear mapping problem with four input-target pairs.
- VI. EXAMPLE OF ANN TRAINING: The network consists of two 3×3 unitary optical interference units with z^2 activations.A constant third input introduces artificial bias terms, and the first output element is used as the prediction.
- VI. EXAMPLE OF ANN TRAINING: After training, the network predictions successfully match the XOR targets in the numerical demonstration.Figure 4 compares mean-squared error across iterations and prediction magnitudes before and after training.
- VI. EXAMPLE OF ANN TRAINING: Training uses a matrix model that provides complex fields at reference points, enabling backpropagation-based gradient computation.The model computes outputs from input modes and phase-shifter settings but is not a first-principles electromagnetic simulation.
- VI. EXAMPLE OF ANN TRAINING: The XOR experiment is presented as a simple demonstration, while the method is stated to apply to more complicated tasks shown in Appendix C.The authors distinguish the demonstration's simplicity from the broader stated applicability of the technique.
VII. DISCUSSION AND CONCLUSION
The discussion frames TRIM as an in-situ, intensity-based implementation of photonic backpropagation that can extend beyond ANNs. Its exactness depends on system assumptions, especially loss and reciprocity conditions.
- VII. DISCUSSION AND CONCLUSION: The method assumes arbitrary complex inputs and parallel integrated intensity detection with virtually no loss.An alternative low-frequency modulation can measure the interference product directly instead of storing isolated intensities.
- VII. DISCUSSION AND CONCLUSION: TRIM is exact for lossless, feed-forward, reciprocal systems, while uniform loss can be accommodated by scaling measured gradients.Mode-dependent loss limits accurate reconstruction of the time-reversed adjoint field, although the simulation retained good accuracy despite significant loss.
- VII. DISCUSSION AND CONCLUSION: The method can handle imperfections such as deviations from perfect 50-50 MZI splits without assuming a specific linear-operation model.Extending the formalism to backscattering requires special treatment for subtracting backscattered contributions.
- VII. DISCUSSION AND CONCLUSION: Dropout offers a photonic route to regularization by temporarily deleting nodes during training and forcing alternative computational paths.The paper describes this procedure as having a strong regularization effect.
- VII. DISCUSSION AND CONCLUSION: TRIM physically propagates the adjoint field and interferes its time-reversed copy with the original field, producing gradients through in-situ intensity measurements.The conclusion also identifies sensitivity analysis and optimization of reconfigurable photonic systems as applications.
- VII. DISCUSSION AND CONCLUSION: The cited related-work passage identifies LeCun, Bengio, and Hinton's Nature article as reference [1].
- VII. DISCUSSION AND CONCLUSION: The approach is intended to make photonic-circuit training efficient and scalable while avoiding brute-force gradient computation or model-based methods.The authors state that such alternatives may not perfectly represent the physical system.
FUNDING INFORMATION
The work acknowledges support from the Gordon and Betty Moore Foundation, the Swiss National Science Foundation, and the Air Force Office of Scientific Research.
- FUNDING INFORMATION: Funding came from the Gordon and Betty Moore Foundation, the Schweizerischer Nationalfonds, and the Air Force Office of Scientific Research.
Appendix A: Non-holomorphic Backpropagation
Appendix A extends photonic backpropagation to non-holomorphic activation functions by separating complex quantities into real and imaginary parts and deriving the corresponding gradients.
- In polar coordinates, the non-holomorphic derivative expression is written element-wise in terms of magnitude and phase variables.
- The derivation begins from the loss gradient with respect to a phase-shifter permittivity in the final optical layer.
- For non-holomorphic activations, the activation and its complex argument are decomposed into real and imaginary components before differentiation.
- The resulting treatment generalizes across layers by using separate definitions for holomorphic and non-holomorphic activation functions.
- Figure 5 illustrates a universal 3 × 3 unitary optical interference unit and the phase-shifter gradient computation associated with Eq. (B4).
Appendix B: Photonic neural network simulation
Appendix B represents photonic neural-network layers with parametrized unitary meshes and shows how phase-shifter gradients can be computed from forward and adjoint fields.
- The simulated optical interference unit is modeled as a product of unitary matrices implemented by Mach–Zehnder-interferometer operations and phase delays.
- For an individual phase shifter, the unitary matrix is split into the portions preceding and following that phase element.
- The phase-shifter gradient is expressed using the field amplitude generated by the forward input and the corresponding adjoint input at the same location.
- Recording amplitudes at all ports during forward and backward propagation enables parallel computation of gradients for every phase shifter.
- The matrix model simplifies the demonstration, whereas the full in situ procedure remains necessary for systems not captured correctly by that model.
Appendix C: Training demonstration
Appendix C demonstrates numerical training of a six-layer photonic network on noisy oblong-ring classification using non-holomorphic activations and gradient-based phase updates.
- The task uses one thousand input–target examples, with labels determined from the magnitude and phase of each two-dimensional input.
- The network contains six 4 × 4 unitary optical layers, |z| activations, a final |z|2 activation, and a softmax output.
- The first two output components determine the predicted class, while cross-entropy cost is evaluated using the target output port.
- Training computes phase gradients for each iteration, sums them across examples, and backpropagates through |z| and |z|2 using the non-holomorphic derivative expression.
- 91% training and test accuracy was achieved after around 4000 iterations, with visual results indicating no overfitting.