Source-linked AI summary

Three-dimensional virtual refocusing of fluorescence microscopy images using deep learning

Yichen Wu, Yair Rivenson, Hongda Wang, Yilin Luo, Eyal Ben-David, Laurent A. Bentolila, Christian Pritz, Aydogan Ozcan

arXiv:1901.11252v2cs.CVcs.LGphysics.app-phphysics.optics

TL;DR

High-throughput 3D fluorescence imaging remains challenging, with illumination-related phototoxicity and bleaching concerns. Deep-Z uses deep neural networks for snapshot 3D refocusing, enhancing depth-of-field by ~20× while correcting drift, tilt, curvature, and optical aberrations.

  • Problem

    High-throughput fluorescence imaging of 3D samples remains challenging, while illumination raises concerns about phototoxicity and fluorescence photobleaching.

  • Method

    Deep-Z is a deep-neural-network framework that provides snapshot 3D refocusing of fluorescence images.

  • Results

    ~20× enhancement of wide-field fluorescence image depth-of-field was demonstrated, alongside correction of sample drift, tilt, curvature, and optical aberrations.

  • Takeaways & Limitations

    Deep-Z digitally expands fluorescence refocusing capability and corrects image aberrations within the demonstrated imaging framework.

  • Takeaways & Limitations

    The retrievable axial range depends on image SNR; when PSF-carried depth information falls below the noise floor, retrieval is limited.

Abstract

from arXiv · show

Three-dimensional (3D) fluorescence microscopy in general requires axial scanning to capture images of a sample at different planes. Here we demonstrate that a deep convolutional neural network can be trained to virtually refocus a 2D fluorescence image onto user-defined 3D surfaces within the sample volume. With this data-driven computational microscopy framework, we imaged the neuron activity of a Caenorhabditis elegans worm in 3D using a time-sequence of fluorescence images acquired at a single focal plane, digitally increasing the depth-of-field of the microscope by 20-fold without any axial scanning, additional hardware, or a trade-off of imaging resolution or speed. Furthermore, we demonstrate that this learning-based approach can correct for sample drift, tilt, and other image aberrations, all digitally performed after the acquisition of a single fluorescence image. This unique framework also cross-connects different imaging modalities to each other, enabling 3D refocusing of a single wide-field fluorescence image to match confocal microscopy images acquired at different sample planes. This deep learning-based 3D image refocusing method might be transformative for imaging and tracking of 3D biological samples, especially over extended periods of time, mitigating photo-toxicity, sample drift, aberration and defocusing related challenges associated with standard 3D fluorescence microscopy techniques.

4 Department of Human Genetics, David Geffen School of Medicine, University of California,

The passage identifies a Los Angeles, California, USA location.

  • The listed location is Los Angeles, California 90095, USA.

6 Department of Genetics, Hebrew University of Jerusalem, Edmond J. Safra Campus, Givat

The passage identifies Jerusalem, Israel, as a listed location.

  • The listed location is Jerusalem, 91904, Israel.

7 Department of Surgery, David Geffen School of Medicine, University of California, Los

Deep-Z addresses limitations of scanned 3D fluorescence microscopy by digitally refocusing single 2D images onto user-defined 3D surfaces. It supports rapid volumetric imaging, aberration correction, and cross-modality matching while retaining standard microscope performance characteristics.

  • 3D fluorescence microscopy commonly relies on axial scanning, which limits acquisition speed and can introduce artifacts from nonsimultaneous measurements.Repeated excitation also raises phototoxicity and photobleaching concerns.
  • Deep-Z uses a trained deep neural network to refocus a single 2D wide-field fluorescence image onto user-defined 3D surfaces without mechanical scanning or additional hardware.The framework is data-driven and does not require a physical imaging-system model or parameter estimation.
  • ~20-fold: Deep-Z expanded an objective’s approximately 1 µm depth-of-field by digitally refocusing a single image across Δz = ±10 µm.The digitally refocused images matched mechanically scanned images over the same axial range.
  • Deep-Z digitally corrected sample drift, tilt, curvature, and spherical aberrations after image acquisition without modifying standard wide-field microscope hardware.The correction used spatially non-uniform depth-modulation patterns to define output surfaces.
  • Deep-Z+ digitally refocused wide-field fluorescence images to match confocal images at corresponding sample planes.This cross-modality framework combines digital refocusing with confocal-like sectioning behavior.
  • Deep-Z output images can improve signal-to-noise ratio relative to corresponding mechanically scanned images by rejecting unlearned noise sources.The reported match between outputs from one measured image and ground-truth stacks from 41 measured images supports the refocusing approach.

3D functional imaging of C. elegans using Deep-Z

Deep-Z generated virtual 3D fluorescence data from single-plane C. elegans videos and enabled tracking of neuron locations and calcium activity across depth.

  • A C. elegans fluorescence video was acquired at a single focal plane using FITC for activity and Texas Red for neuron locations.The video was recorded at approximately 3.6 Hz for approximately 35 seconds.
  • Deep-Z digitally refocused each frame across axial planes from -10 µm to 10 µm, generating a virtual 3D stack for every acquired frame.The axial planes used a 0.5 µm step size.
  • Defocused neurons in the input video were refocused on demand in both fluorescence channels, enabling spatiotemporal tracking of individual neuron activity in 3D.
  • 155 individual neurons were isolated in 3D from Deep-Z output images, with color indicating each neuron’s depth location.
  • Deep-Z output enabled calcium-activity analysis across neurons, including opposing changes in clusters C3 and C2 at t = 14 s.Cluster C3 activity increased while cluster C2 activity decreased at a similar time point.
  • The tracked 3D neuron activity was embedded in a single-plane 2D image sequence and was recovered without mechanical scanning, additional hardware, or sacrificing resolution or imaging speed.

Discussion

Deep-Z enables single-image, user-defined 3D fluorescence refocusing without mechanical scanning or added hardware, while supporting digital correction of aberrations and reduced photodamage. Its refocusing behavior depends on signal-to-noise ratio, training range, and sample-object density.

  • Framework: Deep-Z uses deep neural networks to refocus a single 2D fluorescence image onto user-defined 3D surfaces.The framework does not require a physical imaging-system model, mechanical scanning, additional hardware, or parameter estimation.
  • Results: Deep-Z digitally corrects sample drift, tilt, curvature, and other optical aberrations by applying non-uniform DPMs during inference.These capabilities add degrees of freedom to the imaging system without changing the acquisition hardware.
  • Framework: The network selectively deconvolves spatial features brought into focus while suppressing features that remain out of focus.Refocusing is controlled by a user-defined digital propagation map (DPM).
  • Imaging impact: A single-image acquisition reduces the number of axial planes imaged and therefore helps reduce photodamage, photobleaching, and phototoxicity.This is especially relevant for dynamic or longitudinal biological imaging, where scanning introduces unavoidable time delays.
  • Scope and limitations: Accurate inference becomes challenging when the depth information carried by the point-spread function falls below the image noise floor.Inference was nevertheless robust across a broad range of exposure times not included in training, while fluorescent-object density remains limited as a function of refocusing distance.
  • Results: ~20× enhancement in the depth-of-field of a wide-field fluorescence image was demonstrated using Deep-Z.The reported axial refocusing range is a practical training-data choice rather than an absolute limit.

Fluorescence image acquisition

Fluorescence data were acquired with scanning microscopes, motorized stages, multiple objectives and filter sets, using both image stacks and time-lapsed videos. The resulting images were aligned, normalized, cropped into training pairs, and used to train and validate Deep-Z.

  • Image-stack acquisition: Z-stacks used 0.5 µm or 0.27 µm axial steps, with two channel images acquired at each z-plane.The 20×/0.8NA and 40×/1.3NA objective datasets used the respective step sizes.
  • Time-lapsed acquisition: Dynamic worm recordings used time-lapsed videos with time-multiplexed Texas Red and FITC channels at an average framerate of ~3.6 fps.The maximum camera framerate was 10 fps.
  • Preprocessing: Image stacks were axially aligned, background-subtracted, intensity-normalized, and converted into extended-depth-of-field reference images.Testing without stacks applied the preprocessing steps directly to the input image.

Training and testing of the Deep-Z network

Deep-Z learns to refocus a 2D fluorescence image onto user-defined planes using an appended depth-propagation map. Training combines adversarial and mean-absolute-error objectives, while testing uses only the generator.

  • Network inputs and outputs: The generator accepts a 256 × 256 fluorescence image and a user-defined depth-propagation map, producing the corresponding refocused surface.The input has two channels: fluorescence and DPM; the output represents the fluorescence image at the specified surface.
  • Optimization: Training iteratively minimizes generator and discriminator losses in a conditional generative adversarial framework.The generator loss combines adversarial and MAE terms, while the discriminator evaluates generator outputs against target images.
  • Optimization: The generator loss uses α = 0.02 to regularize the adversarial term with mean absolute error.Adam optimization uses learning rates of 10^-4 for the generator and 3 × 10^-5 for the discriminator.
  • Model selection and testing: The best network is selected by the smallest validation-set MAE, and testing activates only the trained generator.Validation is performed every 50 iterations before blind testing.
  • Computational performance: Training takes approximately 70 hours for 400,000 iterations, while inference takes approximately 0.2 s for 512 × 512 images and 1 s for 1536 × 1536 images.The largest tested field of view was 1536 × 1536 pixels, limited by GPU memory.
  • Image quality evaluation: Output quality is evaluated against ground truth using MSE, RMSE, MAE, correlation coefficient, and SSIM.SSIM assesses structural similarity because conventional error criteria do not fully indicate perceived image similarity.

Tracking and quantification of C. elegans neuron activity

Deep-Z converts time-multiplexed single-plane fluorescence videos into virtual axial stacks for C. elegans neuron analysis. The workflow localizes 155 neurons, quantifies activity, and clusters the 70 most active cells.

  • Virtual 3D reconstruction: Each video frame is registered, appended with DPMs spanning -10 µm to 10 µm in 0.5 µm steps, and passed through Deep-Z to generate a virtual axial stack.The input sequence alternates FITC and Texas Red channels, which are combined into activity and nuclei channels.
  • Virtual 3D reconstruction: Deep-Z generated a virtual axial image stack for each frame, enabling time-resolved three-dimensional neuron measurements.The network was specifically trained for the imaging system used in the experiment.
  • Neuron localization and activity: Watershed segmentation isolated 155 neurons from centroids identified in a median-intensity projection of the red-channel stack.The resulting 3D voxel masks were used for per-neuron activity extraction.
  • Neuron localization and activity: Calcium activity was quantified from the average of the 100 brightest FITC voxels within each neuron mask, then baseline-subtracted as ΔF(t) = F(t) − F0.F0 is the time average of F(t).
  • Activity-pattern clustering: Thresholding selected the 70 most active cells, which were clustered by calcium-activity-pattern similarity using spectral clustering.The eigen-gap heuristic selected k = 3 clusters, followed by k-means clustering of the leading eigenvectors.
  • Cross-modality preparation: Wide-field and confocal image stacks were aligned, stitched, co-registered, and axially centered before paired training examples were created.Four confocal target images were randomly selected for each wide-field input, with DPMs calculated from their centered axial-height differences.
Loading 1901.11252v2…