Source-linked AI summary

NETT: Solving Inverse Problems with Deep Neural Networks

Housen Li, Johannes Schwab, Stephan Antholzer, Markus Haltmeier

arXiv:1803.00092v3math.NAcs.LG

TL;DR

The paper addresses the limited theoretical analysis of neural-network methods for ill-posed inverse problems. It develops NETT, which combines data consistency with a trained neural-network regularizer, and proves well-posedness, convergence, and convergence rates. Numerical tomography results show artifact removal and consistent reconstruction even for image types outside the training data.

  • Problem

    Neural-network methods for inverse problems had few theoretical results despite promising performance, motivating analysis of their stability and convergence.

  • Method

    NETT solves inverse problems by minimizing a data-consistency term together with a neural-network-defined regularizer, supported by absolute Bregman-distance analysis and a proposed training strategy.

  • Results

    NETT is stably solvable with convergence and quantitative error estimates, while tomography experiments remove under-sampling artifacts and preserve resolution beyond the training image type.

  • Takeaways & Limitations

    Learning a regularization functional from one image class can produce good NETT reconstructions for images beyond that class.

  • Takeaways & Limitations

    Detailed investigation of different trained regularizers and comparisons with other methods, computational performance, and real-world data are left for future work.

Abstract

from arXiv · show

Recovering a function or high-dimensional parameter vector from indirect measurements is a central task in various scientific areas. Several methods for solving such inverse problems are well developed and well understood. Recently, novel algorithms using deep learning and neural networks for inverse problems appeared. While still in their infancy, these techniques show astonishing performance for applications like low-dose CT or various sparse data problems. However, there are few theoretical results for deep learning in inverse problems. In this paper, we establish a complete convergence analysis for the proposed NETT (Network Tikhonov) approach to inverse problems. NETT considers data consistent solutions having small value of a regularizer defined by a trained neural network. We derive well-posedness results and quantitative error estimates, and propose a possible strategy for training the regularizer. Our theoretical results and framework are different from any previous work using neural networks for solving inverse problems. A possible data driven regularizer is proposed. Numerical results are presented for a tomographic sparse data problem, which demonstrate good performance of NETT even for unknowns of different type from the training data. To derive the convergence and convergence rates results we introduce a new framework based on the absolute Bregman distance generalizing the standard Bregman distance from the convex to the non-convex case.

1 Introduction

NETT formulates inverse-problem reconstruction as generalized Tikhonov regularization, combining data consistency with a neural-network-defined prior. The paper establishes stable solvability, convergence, quantitative error estimates, and a training strategy, while numerical results demonstrate artifact removal beyond the training distribution.

  • Problem: Inverse problems estimate unknowns from noisy indirect data through a possibly nonlinear operator, requiring regularization when solutions are unstable or underdetermined.The framework covers reflexive Banach spaces and finite-dimensional settings, with noise bounded by a known level δ.
  • NETT formulation: NETT minimizes a data-similarity term plus αR(V, x), where the regularizer is defined by a neural network with trainable parameters.The data term can use squared distance or alternatives such as Kullback-Leibler divergence.
  • Theory: Under reasonable assumptions, NETT is stably solvable and regularized solutions converge to R(V, ·)-minimizing solutions as δ approaches zero.The minimizing solutions satisfy the exact-data constraint F(x)=y0 while minimizing the network regularizer.
  • Theory: The analysis derives norm convergence and convergence rates using absolute Bregman distance, a generalization of standard Bregman distance to non-convex regularization.The framework also introduces total non-linearity for proving strong convergence.
  • Training: A CNN regularizer can assign small values to desirable training phantoms and larger values to artifacts, noise, or undesirable images.The proposed strategy combines encoder-decoder features with training examples representing artifacts and clean images.
  • Comparison: Unlike direct reconstruction networks, NETT combines a neural prior with the forward model, bounding data consistency also for data outside the training set.The paper contrasts this with existing approaches whose non-trivial data-consistency bounds are unavailable beyond training data.
  • Scope: Numerical tomography experiments report artifact removal and preserved high resolution, including phantoms with smooth structures absent from training data.The paper states that detailed comparisons with other methods and real-world data remain outside its scope.

2 NETT regularization

NETT solves ill-posed inverse problems by minimizing a data-consistency term plus a trained neural-network regularizer. Under stated assumptions, the framework provides well-posedness, weak convergence, and strong norm convergence for totally non-linear regularizers.

  • 2.1 The NETT framework: NETT minimizes a data-consistency term combined with a neural-network regularizer encoding prior information about the unknown.The regularizer is built from affine network layers, fixed nonlinearities, and a possibly non-convex functional.
  • 2.1 The NETT framework: The network formulation accommodates Banach-space settings and reduces to standard convolutional neural networks in finite dimensions.Layer spaces can include function spaces, with operations such as pooling or resampling changing the domain between layers.
  • 2.1 The NETT framework: Coercivity is identified as potentially the most restrictive assumption, with skip, residual, layer-wise, leaky-ReLU, and pooling strategies discussed for obtaining it.The theoretical analysis assumes the network regularizer and its parameters are fixed before minimizing the NETT functional.
  • 2.2 Well-posedness and weak convergence: Under Condition 2.2, NETT minimizers exist, are stable under convergent data, and converge weakly to regularizer-minimizing exact solutions.If the minimizing solution is unique, the full sequence converges weakly to it.
  • 2.3 Absolute Bregman distance and total non-linearity: The absolute Bregman distance is introduced to analyze non-convex regularizers, for which the standard Bregman distance may be negative.The paper presents it as a new generalization used to establish strong convergence.
  • 2.3 Absolute Bregman distance and total non-linearity: Total non-linearity of the regularizer strengthens convergence from the weak topology to norm convergence.The strong-convergence theorem applies to minimizing sequences under the stated parameter-choice condition and yields a subsequence, or the full sequence when the minimizer is unique.

3 Convergence rates

The paper derives quantitative NETT error bounds through variational inequalities, including rates in the absolute Bregman distance and extensions to general and non-convex Tikhonov regularizers. These results include exact penalization and recover established sparse-regularization rates in suitable settings.

  • 3.1 General convergence rates result: Under a variational inequality, NETT achieves the error rate E(xα,δ, x+) = O(Φ(τδ)) when α ∼δ/Φ(τδ).The error measure E may be a general functional on the parameter space.
  • 3.1 General convergence rates result: If Φ(t) ≤Ct near zero, exact penalization permits a non-vanishing regularization parameter while βE(xα,δ, x) = O(δ).The stated condition implies that α need not tend to zero as the noise level decreases.
  • 3.2 Rates in the absolute Bregman distance: For Hilbert spaces and bounded linear F, a source condition together with the stated inequality yields rates in the absolute Bregman distance.With squared-norm data similarity, the corresponding conclusions of Theorem 3.1 apply.
  • 3.3 General Tikhonov regularization: The convergence, rate, and absolute-Bregman results extend beyond neural-network regularizers to general Tikhonov regularization under Condition 3.5.The paper notes that items (b)–(d) were previously unavailable for non-convex regularizers.
  • 3.4 Nonlinear ℓq-regularization: For the nonlinear ℓq-type regularizer, choosing α ∼δ gives BR(xα,δ, x+) = O(δ).The construction allows nonlinear functionals and can therefore produce generally non-convex regularizers.
  • 3.4 Nonlinear ℓq-regularization: In sparse settings, the norm rate improves from O(δ1/2) to O(δ1/q) under sparsity and restricted injectivity assumptions.The O(δ1/2) rate reproduces a known classical result, while the improved rate follows from the variational-inequality framework.
  • 3.5 Comparison to the W-Bregman distance: The relation between absolute and W-Bregman distances depends on the particular setting and the chosen family W.The paper identifies analysis using W-Bregman distances as future work.

4 A data driven regularizer for NETT

This section constructs a trained encoder-decoder regularizer for NETT and describes its training and minimization strategy. The regularizer penalizes encoded artifact content while supporting non-convex optimization.

  • 4.1 A trained regularizer: The proposed regularizer uses an encoder-decoder network trained to estimate artifact components from potential solutions.The encoder output norm serves as the trained regularizer.
  • 4.1 A trained regularizer: Training pairs combine back-projection images with artifact residuals and artifact-free examples with zero residuals.For the first examples, r_n = z_n − F+(Fz_n); for later examples, r_n = 0.
  • 4.1 A trained regularizer: The network is optimized so its decoder output approximates each training target using a suitable loss function.Mean absolute error and mean squared error are identified as typical choices.
  • 4.1 A trained regularizer: The trained regularizer is expected to be large for severe artifacts and small for nearly artifact-free images, including some images unlike the training data.The authors report that numerical results confirm applicability beyond the training-image type.
  • 4.2 Minimizing the NETT functional: NETT defines regularized solutions as minimizers of a data-consistency term plus a neural-network regularization term.The resulting optimization problem is non-convex because of the nonlinear network and non-smooth when q = 1.
  • 4.2 Minimizing the NETT functional: An incremental gradient method alternates data-fidelity and regularizer updates to minimize the NETT functional.The authors report favorable performance and stability with respect to tuning parameters, while comparisons with other algorithms remain beyond scope.

5 Application to sparse data tomography

The paper applies NETT to sparse-data photoacoustic tomography, using an encoder-decoder regularizer and iterative minimization to reconstruct images from limited measurements. Experiments report artifact removal, preserved detail, performance on phantoms unlike the training data, and robustness to noise.

  • Application and implementation: NETT is applied to sparse-data photoacoustic tomography using an encoder-decoder network regularizer and an iterative minimization algorithm.The PAT setup recovers initial pressure from acoustic measurements, while the network is implemented through an encoder-decoder scheme.
  • Application and implementation: The sparse PAT system uses 30 equidistant spatial samples, 2,000 temporal samples, α = 1/4, and a zero-image initialization.Reconstructions use 256 × 256 images, a constant step size of 0.4, and Algorithm 1.
  • Results: For the Shepp-Logan phantom, relative L2-errors are 0.262 after 10 iterations and 0.192 after 50, versus 0.338 for FBP.The reported reconstructions remove under-sampling artifacts while preserving high-resolution information.
  • Results: For the blobs phantom, relative reconstruction errors are 0.176 after 10 iterations and 0.102 after 50, versus 0.179 for FBP.This phantom contains smooth structures and differs from the training phantoms; NETT removes artifacts while preserving high resolution.
  • Results: With 5 % additive Gaussian noise, 15-iteration NETT reconstructions achieve relative errors of 0.280 for Shepp-Logan and 0.210 for blobs.The reported noisy reconstructions are free from under-sampling artifacts and retain high-frequency information.
  • Generalization: Across the experiments, a regularizer trained on one phantom class also performs well on images with smooth structures and other types not represented in training.The authors attribute this observed behavior to NETT’s combination of forward-model data consistency and learned regularization.

6 Conclusion and outlook

The paper develops NETT as a theoretically analyzed combination of deep neural networks and Tikhonov regularization for inverse problems. It establishes convergence and well-posedness results and reports initial PAT experiments, while identifying broader comparisons, network designs, algorithms, and applications as future work.

  • Contributions: NETT combines deep neural networks with Tikhonov regularization and provides a framework for solving inverse problems.The regularizer may be user-specified or a CNN trained on appropriate data.
  • Contributions: The analysis establishes well-posedness, weak convergence, norm-convergence, and convergence-rate results for NETT.These results are developed using absolute Bregman distance, extending standard Bregman distance to the non-convex setting.
  • Contributions: Initial sparse-data PAT experiments show that a trained NETT regularizer performs well on phantoms different from the training-data class.The conclusion reports this as an initial numerical demonstration rather than a completed comparative evaluation.
  • Outlook: Future work includes comparisons with other deep-learning and variational methods, alternative minimization algorithms, different network designs, training strategies, and applications to other inverse problems.The authors specifically mention residual or Ivanov regularization, proximal-gradient methods, and semi-smooth Newton methods.
Loading 1803.00092v3…