Source-linked AI summary

A deep convolutional neural network using directional wavelets for low-dose X-ray CT reconstruction

Eunhee Kang, Junhong Min, Jong Chul Ye

arXiv:1610.09736v3cs.CV

TL;DR

Low-dose CT reduces radiation exposure but introduces artifacts that impair image quality, while existing approaches can be computationally expensive or poorly suited to CT-specific noise. The paper addresses this with a directional wavelet-domain CNN using residual learning, and reports effective denoising, faster reconstruction than MBIR, and validation in the 2016 AAPM Low-Dose CT Grand Challenge.

  • Problem

    Low-dose CT introduces severe artifacts that reduce image quality, while model-based denoising is computationally expensive and image-domain methods struggle with CT-specific noise.

  • Method

    The paper uses a deep CNN trained on directional wavelet coefficients, with residual-learning architecture for low-dose CT reconstruction.

  • Results

    The proposed method achieved greater low-dose CT denoising power, faster reconstruction than MBIR, and effectiveness confirmed in the 2016 AAPM Low-Dose CT Grand Challenge.

  • Takeaways & Limitations

    The results support wavelet-domain CNNs as an effective framework for low-dose CT reconstruction.

  • Takeaways & Limitations

    The discussion identifies a scope boundary concerning recent work that goes beyond the present study.

Abstract

from arXiv · show

Due to the potential risk of inducing cancers, radiation dose of X-ray CT should be reduced for routine patient scanning. However, in low-dose X-ray CT, severe artifacts usually occur due to photon starvation, beamhardening, etc, which decrease the reliability of diagnosis. Thus, high quality reconstruction from low-dose X-ray CT data has become one of the important research topics in CT community. Conventional model-based denoising approaches are, however, computationally very expensive, and image domain denoising approaches hardly deal with CT specific noise patterns. To address these issues, we propose an algorithm using a deep convolutional neural network (CNN), which is applied to wavelet transform coefficients of low-dose CT images. Specifically, by using a directional wavelet transform for extracting directional component of artifacts and exploiting the intra- and inter-band correlations, our deep network can effectively suppress CT specific noises. Moreover, our CNN is designed to have various types of residual learning architecture for faster network training and better denoising. Experimental results confirm that the proposed algorithm effectively removes complex noise patterns of CT images, originated from the reduced X-ray dose. In addition, we show that wavelet domain CNN is efficient in removing the noises from low-dose CT compared to an image domain CNN. Our results were rigorously evaluated by several radiologists and won the second place award in 2016 AAPM Low-Dose CT Grand Challenge. To the best of our knowledge, this work is the first deep learning architecture for low-dose CT reconstruction that has been rigorously evaluated and proven for its efficacy.

I. INTRODUCTION

Low-dose CT reduces radiation exposure but introduces degraded image quality and CT-specific artifacts that challenge diagnosis. The paper proposes a wavelet-domain CNN to suppress these artifacts, with experiments showing improved denoising and faster reconstruction than MBIR.

  • Motivation: Reducing X-ray photons lowers radiation exposure but typically degrades image quality because measurements have a low signal-to-noise ratio.
  • Limitations of prior methods: Existing image-domain denoising methods are not ideal for CT-specific noise, including complicated streaking caused by photon starvation and beam hardening.
  • Limitations of prior methods: MBIR models CT physics and noise statistics but typically requires computationally expensive iterative projection/backprojection and trains only a few parameters.
  • Proposed approach: The proposed method applies a deep CNN to directional wavelet coefficients, exploiting directional artifact components and using residual-learning connections.
  • Evaluation: Experiments on the 2016 Low-Dose CT Grand Challenge dataset showed significant improvements over conventional denoising approaches and confirmed the network architecture’s effectiveness.

II. BACKGROUND

Low-dose CT physics produces non-Gaussian artifacts, including streaking from photon starvation and beam hardening. Background denoising methods use image statistics, sparsity, or wavelet coefficients, but are not directly designed for CT artifacts.

  • Low-dose CT physics: X-ray measurements are often modeled with Poisson statistics, and log-transformed sinogram data are often approximated as weighted Gaussian.
  • Low-dose CT physics: Beam hardening violates the linear relation between projection data and attenuation coefficients because lower-energy photons are seldom detected.
  • Low-dose CT physics: Photon starvation occurs when highly attenuating bones absorb many X-ray photons, causing information loss and streaking artifacts along affected directions.
  • Image-domain denoising: Wavelet shrinkage denoises by decomposing images into low- and high-frequency components and thresholding high-frequency coefficients.
  • Image-domain denoising: Total variation denoising is widely used but can produce cartoon-like artifacts.
  • Image-domain denoising: Image denoising methods based on dictionaries, non-local statistics, and BM3D were not designed directly for X-ray CT artifacts.

II.B.2. MBIR for low-dose X-ray CT

MBIR formulates low-dose CT reconstruction as an optimization problem combining data fidelity with regularization, but iterative methods can be computationally expensive. CNNs use layered convolutions and training components such as batch normalization, bypass connections, and contracting paths to improve image processing and training.

  • MBIR formulation: MBIR reconstructs low-dose CT images by minimizing a data-fidelity term plus a regularization penalty.The regularization can impose requirements such as smoothness or sparsity, with λ controlling its strength.
  • MBIR formulation: Common regularizers include total variation, dictionary-based methods, and non-local means approaches.These methods exploit assumptions such as image sparsity under gradient operations or non-local similarity.
  • MBIR limitations: Iterative reconstruction methods have high computational complexity from projection/backprojection operators and iterative optimization for non-differentiable penalties.This limits the computational efficiency of conventional MBIR approaches.
  • CNN foundations: CNNs represent images through successive convolutional and nonlinear functions and learn parameters by minimizing an empirical loss.ReLU is described as a commonly used nonlinear function, while regression denoising can use Euclidean distance loss.
  • CNN foundations: Batch normalization, bypass connections, and contracting paths were developed to improve CNN training and performance capabilities.Bypass connections preserve image details and support gradient back-propagation, while contracting paths retain high-resolution features after down-sampling.

III. METHOD

The proposed network is motivated by combining directional wavelet decomposition with deep CNN denoising for complex low-dose CT noise. Its design also uses large training data and residual learning to capture diverse information and simplify learning.

  • Motivation: A directional wavelet transform such as a contourlet can decompose directional noise components to facilitate deep-network training.The transform is intended to separate directional structure in the noise before CNN processing.
  • Motivation: Low-dose CT contains complex noise, while CNNs have potential to remove such noise.This pairing motivates applying deep learning to low-dose CT denoising.
  • Motivation: Deep neural networks can capture various types of information from large amounts of training data.This observation supports the use of a deep architecture for the reconstruction task.

III.A. Contourlet transform

The method uses a non-subsampled contourlet transform to represent low-dose CT images across multiple scales and directions. A wavelet-domain CNN then denoises the resulting coefficients using residual connections and correlations across subbands.

  • Contourlet transform: The contourlet transform combines multiscale decomposition with directional decomposition of high-pass subbands.High-pass and low-pass subbands are formed first, and directional filter banks divide high-pass content into directional subbands.
  • Contourlet transform: The non-subsampled contourlet transform is shift-invariant because its filter banks use neither down-sampling nor up-sampling.This representation uses non-subsampled pyramids and non-subsampled directional filter banks.
  • Contourlet transform: The transform produces four levels with 8, 4, 2, and 1 directional subbands, respectively.Together these levels generate 15 channels for the network input.
  • Contourlet transform: Low-dose CT edges and noise appear in high-frequency components, including streaking noise between bones.This concentration motivates denoising the directional high-frequency bands.
  • Wavelet-domain network: The proposed CNN denoises wavelet coefficients while exploiting inter- and intra-scale correlations through a trainable shrinkage operator.The noisy image is first decomposed into contourlet coefficients before network processing.
  • Wavelet-domain network: Residual and bypass connections pass low-frequency coefficients directly and ease deep-network training.Channel concatenation provides multiple gradient paths, enabling faster end-to-end training and better denoising performance.

III.C. Network training

Network training uses mini-batch stochastic gradient descent with regularization, data augmentation, and gradient clipping. The reported settings support rapid convergence while maintaining performance across a tested range of regularization strengths.

  • Optimization: Training minimizes a loss function with an additional L2 regularization term using mini-batch stochastic gradient descent.The convolution weights are initialized from random Gaussian distributions.
  • Hyper-parameter analysis: The regularization parameter λ varied within [10^-5, 10^-3], and network performance was not sensitive to that choice.Similar performance was observed across the tested range.
  • Optimization: Gradient clipping to [−10^-3, 10^-3] prevents exploding gradients during high-learning-rate training.The authors report that this setting facilitates rapid convergence and avoids gradient explosion.
  • Training data: The mini-batch size was ten, using randomly selected 55×55×15 wavelet-coefficient blocks.Training CT images were also randomly flipped or rotated for data augmentation.
  • Implementation: The proposed network was implemented with MatConvNet on MATLAB, with training environments reported in the accompanying tables.The supplied table captions identify network hyper-parameters and training-data specifications.

III.D. Data Set

The study generated CT training and test images from projection data, using normal-dose reconstructions as ground truth and quarter-dose images as noisy inputs. Test data comprised quarter-dose images from 20 patients, whose results were evaluated by radiologists.

  • Training data: The training data contained 3642 slices from 2304 projection views, with training performed using randomly extracted subsets because memory was insufficient for the entire set.Two hundred slices were sampled and changed every 50 epochs.
  • Test data: The Challenge test set contained only quarter-dose exposure images, comprising 2101 slices from 20 patients, and radiologists evaluated the resulting reconstructions.The paper reports representative test-data reconstructions.
  • Data generation: Projection data were rebinned into conventional fanbeam data and reconstructed as 512×512 CT images using filtered backprojection.The rebinned slices had 1 mm thickness.
  • Training pairs: Normal-dose reconstructions served as ground truth, while quarter-dose reconstructions supplied the noisy inputs for supervised learning.The network learned the mapping between these image types.
  • Evaluation preparation: For the Grand Challenge submission, 1 mm images were averaged across three adjacent slices to form 3 mm images for evaluation.The network was also retrained directly with 3 mm averaged CT images.

III.E. Image Metrics

Image quality was assessed against normal-dose ground truth using PSNR and NRMSE, alongside visual and profile-based evaluations. The proposed method reduced noise and artifacts while preserving anatomical edges and details, with approximately 1.6 seconds required per 512×512 slice.

  • Quantitative metrics: PSNR and NRMSE were calculated by comparing denoised quarter-dose reconstructions with normal-dose ground-truth images.The metrics are derived from MSE, with PSNR based on the maximum ground-truth intensity and NRMSE using RMSE and the ground-truth intensity range.
  • Profile analysis: Profile evaluations showed that the proposed network reduced noise and streaking artifacts while retaining peak points in both training and test data.The test-data profiles also demonstrated feasible noise reduction.
  • Limitations: Some results appeared blurred and lost high-frequency textures because the 1 mm normal-dose ground-truth images still contained noise.This noise reduced the accuracy of supervised learning.
  • Computational cost: 1.6 seconds per 512×512 CT slice was the average MATLAB processing time on a dual-GPU system.Whole-body CT data could be processed in approximately 2.1–3.3 minutes.
  • Anatomical detail: The reconstructions preserved liver textures and lesion-related details, supporting clearer determination of lesion locations in transverse, coronal, and sagittal views.Lesions and liver vessels were highlighted in the test examples.
  • Image quality: The proposed network reduced noise across a wide range of levels while maintaining edge information in reconstructed CT images.Visual examples show clearer organ boundaries and details.

IV.B. Role of the wavelet transform

The wavelet transform was evaluated by comparing the proposed wavelet-domain CNN with an otherwise identical image-domain CNN. The wavelet-domain model outperformed the image-domain baseline on PSNR and NRMSE and recovered more structures visually.

  • Comparison design: The baseline CNN matched the proposed architecture but used image patches as inputs and outputs instead of local wavelet coefficients.Both networks were trained on data from nine patients and evaluated on another patient.
  • Quantitative comparison: The wavelet-based CNN outperformed the image-domain CNN in both PSNR and NRMSE.The comparison was reported in Fig. 12.
  • Visual comparison: Differences between the methods were concentrated mainly around image edges, and the image-domain CNN failed to recover many structures in magnified regions.The unrecovered structures were identified with red arrows.

IV.C. Analysis of residual learning techniques

Residual learning was tested through bypass connections and baseline comparisons, with the proposed architecture shown to improve training. The method’s broader scope remains constrained by dose-specific noise and texture differences from MBIR reconstructions.

  • Residual architecture: The proposed residual architecture used an external low-frequency bypass and internal module bypass connections, whereas the baseline omitted these connections.The network also combined module outputs in a final concatenated layer.
  • Residual-learning analysis: The residual-learning comparison confirmed that the architecture was helpful for training the proposed network.The study compared the proposed model with an identical architecture without residual learning.
  • Data-driven scope: The authors argued that the method can learn organ-, protocol-, and hardware-dependent noise from larger training data sets.This was presented as an advantage over MBIR approaches, which usually train a single regularization parameter.
  • Dose-level scope: The method was designed for quarter-dose data, and applying it to lower-dose noise produced blurring artifacts.Different noise levels require additional training data.
  • Projection-domain scope: Projection-domain application was outside the study’s scope because projection data are related across angles, complicating patch processing.A large-receptive-field network was suggested as a possible way to address this difficulty.
  • Texture limitation: Radiologists reported that the deep-learning reconstruction texture differed from MBIR textures, which are also important diagnostic features.The authors identified this as a follow-up problem requiring a new deep-learning method.

VI. CONCLUSION

The paper introduces a deep CNN framework for low-dose CT reconstruction that combines directional wavelets with deep convolutional learning. It reports stronger denoising, faster reconstruction than MBIR methods, and rigorous validation in the 2016 AAPM Low-Dose CT Grand Challenge.

  • The proposed framework combines a deep convolutional neural network with a directional wavelet approach for low-dose CT reconstruction.
  • The method provides greater de-noising power for low-dose CT and reconstructs images much faster than MBIR methods.
  • The proposed network’s effectiveness was confirmed in the 2016 AAPM Low-Dose CT Grand Challenge.
  • The authors present the method as an innovative framework for low-dose CT research.
Loading 1610.09736v3…