Source-linked AI summary

Deep learning in ultrasound imaging

Ruud JG van Sloun, Regev Cohen, Yonina C Eldar

arXiv:1907.02994v2eess.SPcs.LGeess.IV

TL;DR

Ultrasound receive processing faces increasing data, computational, and imaging-quality demands across beamforming, Doppler, clutter suppression, and super-resolution. This paper develops model-based deep learning and artificial-agent approaches that exploit signal structure, reporting improved image quality, faster adaptive beamforming, and enhanced clutter separation. These results support modular, application-adaptive ultrasound processing chains while retaining important computational and modeling constraints.

  • Problem

    Ultrasound processing requires methods that handle demanding receive-chain tasks while maintaining image quality, and existing model-based methods can have high computational complexity or constrained operating assumptions.

  • Method

    The paper combines deep learning with signal models and structural priors to build artificial agents, learned processors, and deep-unfolded approximations for ultrasound reconstruction tasks.

  • Results

    Deep learning approaches improve beamforming resolution and contrast, reduce adaptive beamforming computation by more than 400×, and improve clutter separation over SVD filtering.

  • Takeaways & Limitations

    Modular learned components can be combined into fully adaptive ultrasound imaging chains dedicated to particular applications.

Abstract

from arXiv · show

We consider deep learning strategies in ultrasound systems, from the front-end to advanced applications. Our goal is to provide the reader with a broad understanding of the possible impact of deep learning methodologies on many aspects of ultrasound imaging. In particular, we discuss methods that lie at the interface of signal acquisition and machine learning, exploiting both data structure (e.g. sparsity in some domain) and data dimensionality (big data) already at the raw radio-frequency channel stage. As some examples, we outline efficient and effective deep learning solutions for adaptive beamforming and adaptive spectral Doppler through artificial agents, learn compressive encodings for color Doppler, and provide a framework for structured signal recovery by learning fast approximations of iterative minimization problems, with applications to clutter suppression and super-resolution ultrasound. These emerging technologies may have considerable impact on ultrasound imaging, showing promise across key components in the receive processing chain.

I. INTRODUCTION

Ultrasound’s expanding use of compact, ultrafast, and advanced imaging systems increases receive-processing demands, especially from high data rates and unfocused transmissions. The paper examines deep learning and signal-structure methods across the imaging chain, including compressive sampling and adaptive processing.

  • Ultrasound combines cost-effectiveness, portability, real-time interaction, and broad medical use, supporting point-of-care imaging across diverse settings.
  • Compact probes, 3D imaging, high-frame-rate schemes, and advanced applications increase data rates and burden probe communication and reconstruction algorithms.
  • The paper advocates deep learning methods that exploit signal structure and data to develop robust, data-efficient solutions across key ultrasound processing components.
  • Ultrafast unfocused transmissions create hardware and algorithmic burdens, requiring advanced receive beamforming and clutter suppression for satisfactory image quality.
  • Modern receive processing uses digital parallel beamforming and coherent compounding, but dense matrix probes make the required coaxial connections infeasible.
  • Sub-Nyquist sampling under a finite-rate-of-innovation model can reduce sampling rates by up to 28 fold.

C. B-mode, M-mode, and Doppler

Ultrasound supports anatomical, flow, tissue-motion, and microvascular imaging, but extracting functional signals and resolving microvasculature remains challenging. The paper applies structure-aware deep learning to processing tasks including clutter suppression and super-resolution.

  • B-mode imaging forms anatomical images by envelope-detecting beamformed signals and applying dynamic-range compression before scan conversion.
  • Doppler processing estimates blood-flow or tissue-displacement velocities through Color Doppler and Spectral Doppler methods.
  • Ultrasound advanced applications include elastography, which measures mechanical tissue parameters from displacement or shear-wave propagation.
  • Low-velocity microvascular flow is difficult to detect because its Doppler spectrum overlaps strong tissue clutter; CEUS addresses microvascular visualization with gas-filled microbubbles.
  • Ultrasound localization microscopy depends on detecting, isolating, and localizing microbubbles using tuned clutter suppression and concentration constraints.
  • The paper develops model-based, data-driven, and structure-aware deep learning processors that act across the imaging chain as adaptive agents or signal processors.

A. Beamforming

Beamforming methods range from conventional delay-and-sum to learned and model-constrained adaptive approaches. Deep learning can improve reconstruction quality while replacing computational bottlenecks in adaptive beamforming with compact estimators.

  • Delay-and-sum beamforming is widely used for real-time reconstruction because of its low complexity, but its fixed design can deteriorate image quality.
  • 1) Deep neural networks as beamformers:: Deep learning beamformers include learned delay layers, stacked autoencoders, encoder-decoder networks, and fully convolutional mappings from channel data to beamformed outputs.
  • Very deep beamforming networks may require vast RF channel datasets and have memory footprints that complicate resource-limited implementations.
  • 2) Leveraging model-based algorithms:: MVDR adaptively optimizes apodization weights from signal statistics but requires covariance-matrix inversion and becomes impractical for typical large ultrasound arrays.
  • 2) Leveraging model-based algorithms:: A neural artificial agent can replace the MVDR weight-estimation bottleneck by predicting per-pixel apodization weights under a close-to-distortionless constraint.
  • 2) Leveraging model-based algorithms:: For 128 elements, learned adaptive apodization requires 74656 FLOPS versus >2,097,152 FLOPS for MVDR, yielding more than 400× speed-up in reconstruction time.
  • 1) Deep neural networks as beamformers:: Adaptive beamforming by deep learning provides reduced clutter, enhanced tissue contrast, and improved axial and lateral resolution over standard delay-and-sum beamforming.Reported resolution values are 0.43 mm versus 0.34 mm axially and 0.85 mm versus 0.70 mm laterally; contrast-to-noise ratio is 10.96 dB versus 11.48 dB.

3) Design and training considerations:

Ultrasound neural networks require design choices that accommodate the large dynamic range and later transformations of channel and beamformed signals. The paper uses dynamic-range-preserving activations and logarithmic losses to support learning for ultrasound processing.

  • Activation functions: Large dynamic range in radio-frequency ultrasound data makes standard ReLU activation behavior difficult to maintain during training.The signals may not remain in the activation range where gradients are sufficiently large.
  • Activation functions: The anti-rectifier preserves positive and negative responses while providing the nonlinearity needed to learn complex representations.It avoids vanishing gradients and dying nodes for negative values.
  • Activation functions: The anti-rectifier is suited to radio-frequency or IQ-demodulated ultrasound channel data because it preserves signal dynamic range.The scheme is used for the reported results.
  • Training loss: Beamforming training incorporates envelope detection and logarithmic dynamic-range compression through a mean squared logarithmic error.The loss transforms beamforming errors to reflect subsequent processing of envelope-detected signals.
  • Training loss: The logarithmic loss compares predicted beamformed responses with target signals, including MVDR outputs and responses after neural-network-derived apodization.The prediction vector contains responses for all pixels, while the target vector contains the corresponding beamformed signals.

B. Adaptive spectral estimation for spectral Doppler

Adaptive spectral estimation improves the time-frequency resolution tradeoff in spectral Doppler but is computationally expensive because it requires covariance-matrix inversion. The paper uses neural networks as artificial agents to estimate filterbank coefficients efficiently, while addressing training and sampling considerations.

  • Motivation: Spectral Doppler estimates blood and tissue velocity distributions from slow-time pulse-echo snapshots, traditionally using Fourier-transform periodograms such as Welch’s method.High spectral resolution with these methods requires long coherent processing intervals.
  • Adaptive estimation: Data-adaptive spectral estimators provide improved spectral estimates and resolution for a given temporal resolution by using content-matched filterbanks.Temporal resolution is determined by the coherent processing interval, which depends on pulse repetition frequency and the number of slow-time snapshots.
  • Adaptive estimation: Adaptive spectral estimation lowers the required observation window and increases spectral fidelity, but covariance-matrix inversion creates high computational complexity.The input covariance matrix is formed from the slow-time signal vector.
  • Artificial-agent implementation: A neural network takes beamformed slow-time RF data and outputs filter coefficients for each filter in the spectral-estimation filterbank.The filterbank processes the slow-time input signal to produce a spectral estimate.
  • Artificial-agent implementation: The artificial agent is trained with mean squared logarithmic error against a high-quality adaptive Capon spectral estimator and uses 128 subnetworks for 128 filters.Each four-layer fully connected subnetwork predicts the coefficients of one filter.
  • Training considerations: Training uses non-saturating activations for large-dynamic-range inputs, log-transformed loss for decibel spectra, and a penalty on filterbanks deviating from unity frequency response.These choices are presented as training considerations for the artificial agent.
  • Sampling considerations: Extensions are described for periodically gapped and nested slow-time sampling because spectral Doppler is interleaved with B-mode imaging in Duplex mode.These extensions target estimators that can handle gaps or sparse sampling.

C. Compressive encodings for tissue Doppler

Deep encoder-decoder and algorithm-unfolding methods compress Doppler data and approximate structured optimization for ultrasound signal extraction and clutter suppression. These approaches reduce communication or computational demands while preserving or improving selected signal-quality measures.

  • Motivation and compressive acquisition: Limited cable bandwidth, especially in catheter transducers, motivates reduced-rate sampling and neural encoders for ultrasound channel data.Compressive sub-Nyquist sampling can reduce acquisition rates, while probe-side encoders may further alleviate probe-scanner communication.
  • Compressive tissue Doppler: Signal-extracting encoder-decoders decode tissue Doppler or velocity information from compressed IQ-demodulated input rather than reconstructing the full input signal.The network is trained to preserve the functionality and performance of a conventional Kasai autocorrelator using full uncompressed IQ data.
  • Compressive tissue Doppler: IQ compression rates as high as 32 retain reasonable Doppler quality, with relative phase RMSE of approximately 0.02.Lower compression rates reduce the error, whereas higher compression rates increase spatial consistency and suppress unrepresentable spurious variations.
  • Unfolding robust PCA: Clutter suppression can be formulated as low-rank-and-sparse decomposition and solved with an algorithm-unfolded network whose layers embed soft-thresholding and singular-value thresholding.CORONA is a CNN tailored to solving RPCA while explicitly incorporating the signal structure of the optimization problem.
  • Unfolding robust PCA: FISTA improves clutter suppression over SVD filtering but requires tuned thresholds and may need many iterations, motivating fixed-complexity networks with automatic parameter adjustment.The need for automatic threshold adaptation is linked to real-time imaging constraints.
  • Unfolding robust PCA: CORONA outperforms SVD filtering and FISTA on contrast-enhanced ultrasound scans, with CR approximately 15 dB versus approximately 4.6 dB for FISTA and approximately 5.4 dB for SVD filtering.In most cases, CORONA’s performance is about an order of magnitude better than SVD, with only a moderate increase in complexity.

A. Ultrasound localization microscopy

Ultrasound localization microscopy circumvents diffraction-limited resolution by localizing isolated microbubble point sources on a sub-diffraction grid. Its fidelity depends on localization accuracy and the number of localized bubbles, creating a sparsity-versus-acquisition-time trade-off.

  • Diffraction limit: Ultrasound resolution is fundamentally limited by diffraction, while increasing transmit frequency improves wavelength but reduces penetration depth.The minimum distance between separable scatters is described as half a wavelength.
  • ULM principle: ULM achieves super-resolution by isolating individual point sources and precisely localizing their centers on a sub-diffraction grid.The method adapts a principle from super-resolution fluorescence microscopy to ultrasound imaging.
  • Acquisition trade-off: ULM fidelity depends on the number of localized microbubbles and localization accuracy, requiring sparse microbubble solutions that can lead to tediously long acquisitions.Diluted microbubble concentrations are typically used to isolate backscattered echoes on regular ultrasound systems.

B. Exploiting signal structure

Ultrasound localization microscopy can exploit sparsity and temporal structure to recover high-resolution microbubble distributions, improving localization at high concentrations. However, iterative sparse recovery is slow, parameter-sensitive, and based on an approximate linear measurement model.

  • Frame-level signal structure: Sparse recovery models each contrast-enhanced frame as measurements of a sparse microbubble distribution convolved with shifted point-spread functions and corrupted by noise.The high-resolution sparse vector describes microbubble locations, while the measurement matrix contains shifted point-spread functions.
  • Frame-level signal structure: An ℓ1-regularized inverse problem promotes sparse high-resolution reconstructions by penalizing the number and magnitude of non-zero entries.The regularization parameter λ controls the influence of the ℓ1 penalty.
  • Temporal signal structure: FISTA solves the sparse-recovery problem frame by frame, after which estimated microbubble distributions are summed across frames to form the super-resolution image.Temporal models can additionally exploit sparsity in a correlation domain and microbubble motion.
  • Benefits and limitations: Sparse recovery improves localization precision and recall at high microbubble concentrations, but proximal-gradient methods typically require many iterations and careful parameter tuning.The linear measurement model also approximates a nonlinear relation and becomes less adequate for closely spaced microbubbles with strong RF interference.

C. Deep learning for fast high-fidelity sparse recovery

Deep-ULM uses a convolutional encoder-decoder trained on realistic ultrasound simulations to rapidly reconstruct sparse, high-resolution microbubble images from low-resolution frames. It achieves substantially improved resolution and speed while retaining a data-driven reconstruction workflow.

  • Encoder-decoder architecture: Deep-ULM maps low-resolution contrast-enhanced ultrasound frames to sparse localizations on an 8 times finer grid using a convolutional neural network.The network solves the inverse problem using simulations of the corresponding ultrasound acquisition.
  • Encoder-decoder architecture: The fully convolutional encoder-decoder compresses input frames into a latent representation before decoding them into high-resolution sparse outputs, with simultaneous denoising through the latent space.Its encoder uses contracting convolutional blocks, while the decoder upsamples toward the sparse output.
  • Training strategy: The simulations incorporate system point-spread-function estimates, RF modulation frequency, pixel spacing, and noise, clutter, and artifacts sampled from real measurements.These choices expose the network to acquisition-specific signal and artifact statistics during training.
  • Training strategy: The training loss combines a sparsity penalty on reconstructions with Gaussian-smoothed targets to penalize both insufficient sparsity and localization errors.The relative weighting of the sparsity term is controlled by γ.
  • Performance: 20-30 µm resolution represents a 4-5 fold improvement over standard imaging with the adopted linear 15-MHz transducer.The reconstruction qualitatively shows higher resolution and contrast than diffraction-limited maximum intensity projection.
  • Performance: 100 milliseconds per frame on a 4096 × 1328 grid makes deep-ULM about four orders of magnitude faster than Fourier-domain FISTA sparse recovery.The reported timing uses GPU acceleration.

2) Deep unfolding for robust and fast sparse decoding:

Deep unfolded ULM embeds the structure of ISTA into a compact trainable network, replacing repeated optimization with a fixed-length learned scheme. Both deep-learning methods outperform conventional approaches on high-concentration simulations, while deep unfolded ULM generalizes better to in-vivo acquisitions and uses far fewer computational resources.

  • Architecture: Deep unfolded ULM incorporates microbubble sparsity into a compact architecture inspired by proximal-gradient methods.Its design targets robustness by exploiting the underlying sparse decoding structure.
  • Learned iterative scheme: Unfolding ISTA produces a K-layer feedforward network with trainable convolutional kernels and a learned shrinkage parameter at each iteration.A smooth sigmoid-based thresholding operation replaces the proximal soft-thresholding operator to avoid vanishing gradients in its dead zone.
  • Results: Both deep-learning methods significantly outperform standard ULM and FISTA sparse decoding for high microbubble concentrations on synthetic data.Deep-ULM has higher recall and lower localization errors than deep unfolded ULM in these simulations.
  • Results: Deep unfolded ULM produces higher-fidelity super-resolution images on in-vivo ultrasound data and translates better to real acquisitions than the larger encoder-decoder network.The simulation ranking reverses for qualitative in-vivo fidelity.
  • Efficiency: 506 trainable parameters and just over 1000 FLOPS make deep unfolded ULM substantially more efficient than the encoder-decoder, which has almost 700000 parameters and over 4 million FLOPS.The compact model also has lower memory footprint, reduced power consumption, and higher inference rates.
  • Limitation: Strong bone reflections remain more prominent with the compact unfolding scheme, despite its improved robustness and real-data generalization.This artifact is visible in the bottom-left region of the rat spinal-cord reconstruction.

V. OTHER APPLICATIONS OF DEEP LEARNING IN ULTRASOUND

Deep learning in ultrasound extends beyond receive processing to automated image analysis and tissue characterization. Reported applications include scan-plane detection, lesion and artifact identification, anatomical segmentation, registration, and data-driven elasticity imaging.

  • Computer vision: Computer-vision methods apply deep learning to ultrasound images to accelerate and potentially improve clinical diagnostics.The paper distinguishes these applications from its primary focus on ultrasound-specific receive processing.
  • Computer vision: Deep learning can simplify prenatal screening by enabling real-time detection and localization of relevant scan planes and structures.Such examinations otherwise require substantial training to identify the planes and structures of interest efficiently.
  • Computer vision: Other image-analysis applications include thyroid nodule detection, breast-tumor identification and segmentation, lung B-line localization, and real-time TRUS anatomical segmentation.Anatomical landmarks and boundaries can also support voxel-level registration.
  • Tissue characterization: Data-driven elasticity imaging uses neural networks to estimate spatially varying linear elastic material properties from force-displacement measurements without prior constitutive-model or material-property assumptions.This represents a learning-based approach to extracting tissue parameters rather than analyzing images alone.

VI. DISCUSSION AND FUTURE PERSPECTIVES

The paper presents deep learning as a collection of adaptive, structure-aware building blocks for ultrasound processing, while identifying data, hardware, and deployment challenges. It envisions integrated systems combining intelligent probes, edge or cloud computation, and application-specific imaging.

  • Scientific opportunities: Deep learning methods benefit from incorporating ultrasound signal priors and structure, including through deep unfolding and learned beamforming.The paper specifically highlights clutter suppression, super-resolution imaging, and learned beamforming as examples.
  • System integration: Independent neural processors and artificial agents can operate on channel data, IQ data, or images, and may be optimized together as adaptive imaging chains.The proposed building blocks target distinct applications but can form a holistic processing chain dedicated to a specific application.
  • Challenges: Real-time channel-data processing remains challenging because radio-frequency channel data has a large dynamic range and differs from typical image-analysis inputs.The paper discusses concatenated rectified linear units as a possible alternative to common activation functions.
  • System integration: Sub-Nyquist sampling could reduce transmitted channel-data rates, enabling wireless probes to send low-rate data to remote or cloud processors for deep learning.The paper presents this as a route toward intelligent image formation and advanced ultrasound processing.
  • Challenges: Deep learning still requires substantial training data, despite approaches discussed for improving data efficiency and robustness.In supervised learning, the required targets depend on the application and goal.
  • Deployment: Deployment may use GPUs, FPGAs, ASICs, NPUs, or TPUs depending on whether systems prioritize high-end performance, low power, or edge inference.The hardware options span remote processors, resource-limited settings, and consumer devices.
  • Deployment: Continuous lifetime learning in deployed ultrasound systems increases the relevance of unsupervised and self-supervised learning.The paper connects this need to artificial agents adapting to the large amount of data available after deployment.
  • Future perspectives: The paper envisions smart wireless probes connected to cloud computation, with application-specific AI imaging modes and continuously learning devices.This proposed architecture is associated with more portable, cost-effective, and intelligent ultrasound imaging.
Loading 1907.02994v2…