Source-linked AI summary

Learning Proximal Operators: Using Denoising Networks for Regularizing Inverse Imaging Problems

Tim Meinhardt, Michael Moeller, Caner Hazirbas, Daniel Cremers

arXiv:1704.03488v2cs.CV

TL;DR

The paper addresses the costly retraining required by learning-based inverse imaging methods when tasks or operators change. It replaces the proximal operator in convex optimization with a fixed denoising network, and finds state-of-the-art results on deblurring and demosaicking while analyzing algorithmic fixed points and parameter effects.

  • Problem

    Learning-based inverse imaging methods often require costly retraining when the operator, task, noise, or fidelity setting changes.

  • Method

    The paper uses a learned denoising network as the proximal operator in convex optimization algorithms, combining it with independently chosen data fidelity terms and explicit priors.

  • Results

    A fixed denoising network achieves state-of-the-art reconstruction results on exemplary deblurring and demosaicking problems.

  • Takeaways & Limitations

    The approach can reduce problem-specific retraining and use learned natural-image priors across inverse imaging tasks.

  • Takeaways & Limitations

    Convergence guarantees for the four algorithmic schemes are limited to sufficiently friendly convex functions or specific additional assumptions in some nonconvex settings.

Abstract

from arXiv · show

While variational methods have been among the most powerful tools for solving linear inverse problems in imaging, deep (convolutional) neural networks have recently taken the lead in many challenging benchmarks. A remaining drawback of deep learning approaches is their requirement for an expensive retraining whenever the specific problem, the noise level, noise type, or desired measure of fidelity changes. On the contrary, variational methods have a plug-and-play nature as they usually consist of separate data fidelity and regularization terms. In this paper we study the possibility of replacing the proximal operator of the regularization used in many convex energy minimization algorithms by a denoising neural network. The latter therefore serves as an implicit natural image prior, while the data term can still be chosen independently. Using a fixed denoising neural network in exemplary problems of image deconvolution with different blur kernels and image demosaicking, we obtain state-of-the-art reconstruction results. These indicate the high generalizability of our approach and a reduction of the need for problem-specific training. Additionally, we discuss novel results on the analysis of possible optimization algorithms to incorporate the network into, as well as the choices of algorithm parameters and their relation to the noise level the neural network is trained on.

1. Introduction

The paper replaces the proximal operator in convex optimization algorithms with a fixed learned denoising network, allowing the data fidelity term to change across inverse imaging tasks without retraining. It analyzes the resulting algorithmic schemes and reports state-of-the-art performance close to problem-specific networks.

  • Motivation: Linear inverse imaging problems estimate an unobserved image u from measurements f related through a linear operator A and noise.The underlying continuous problem is ill-posed and solutions can be highly sensitive to input data.
  • Proposed approach: A fixed denoising CNN can replace the proximal operator, while changing the reconstruction task changes only the data fidelity term.The same network can therefore be used for tasks such as deblurring and demosaicking without retraining.
  • Variational formulation: Variational reconstruction combines a data fidelity measure H_f(Au) with a regularizer R that encodes prior information about the expected solution.The fidelity term relates measurements to the estimate, while regularization stabilizes reconstruction.
  • Motivation: Deep networks perform well on inverse problems but require costly training when the operator A changes, motivating a more reusable approach.Training also depends on sufficient data and network-architecture expertise.
  • Contributions: Using a fixed denoising network as the PDHG proximal operator yields state-of-the-art results close to problem-specific networks.The paper also analyzes multiple optimization schemes and their fixed points, plus the effects of step size and denoising strength.

2. Related work

Related work established plug-and-play proximal methods using classical denoisers such as TV-inspired methods, NLM, and BM3D. This paper extends that direction to deep convolutional denoising networks and analyzes their behavior theoretically and numerically.

  • Variational methods: Classical variational methods use regularizers such as total variation to suppress noise while preserving image discontinuities.TV penalizes the image-gradient norm.
  • Plug-and-play methods: Convex optimization algorithms often require only evaluations of the regularizer’s proximal operator, motivating its replacement by a denoising method.Prior work used NLM and BM3D as customized proximal operators.
  • Prior work: Customized proximal operators have been applied to several imaging problems, but earlier work focused mainly on patch-based denoising methods.Examples include Poisson denoising, tomography, super-resolution, and hyperspectral image sharpening.
  • Novelty: The paper proposes deep convolutional denoising networks as proximal operators and studies their behavior numerically and theoretically.This extends customized-proximal approaches beyond patch-based denoisers.

3. Learned proximal operators

The paper replaces regularization proximal operators in convex optimization schemes with learned denoising networks, while retaining independently specified data fidelity terms. It analyzes algorithmic equivalence and parameter choices for incorporating these networks.

  • 3.1. Motivation via MAP estimates: MAP formulations make regularization difficult to hand-design for complex natural images, motivating learned image priors while leaving the data term tied to the forward model and noise model.The regularizer corresponds to the negative log-probability of an image, whereas Gaussian noise yields a squared ℓ2 data fidelity term.
  • 3.2. Algorithms for learned proximal operators: The proposed algorithmic scheme replaces the regularization proximal operator in methods such as PG, ADMM, and PDHG with a denoising network G.PDHG variants can decouple linear operators in the regularization from the remaining proximity computation.
  • 3.2. Algorithms for learned proximal operators: For an arbitrary continuous G, the fixed-point equations of the resulting PG, ADMM, PDHG1, and PDHG2 schemes are equivalent.The paper investigates fixed points because convergence guarantees generally require sufficiently friendly convex functions or additional assumptions in nonconvex settings.
  • 3.3. Parameters for learned proximal operators: Any γ > 0 is equivalent to γ = 1 with a newly weighted data fidelity term, so changing γ merely changes the data fidelity parameter.This allows denoising strength to remain fixed while smoothness is controlled through data fidelity, avoiding a difficult-to-tune internal step-size hyperparameter.
  • 3.3. Parameters for learned proximal operators: The optimal data fidelity parameter α is well approximated by a parabola in the denoiser training noise level σ, consistent with the relation α = p σ2.The result comes from exhaustive α selection maximizing PSNR for each network in the deconvolution experiment.
  • 3.3. Parameters for learned proximal operators: Sufficiently large denoiser training noise levels give good results across a broad range, whereas networks trained on very small noise do not.Rescaling regularization and data fidelity does not preserve the learned-network results at each tested data point.

4. Numerical implementation

The method integrates a fixed deep denoising network into PDHG as an implicit prior, while allowing explicit regularizations and data fidelity terms to be combined through prior stacking.

  • Algorithmic framework and prior stacking: PDHG replaces a regularization proximal operator with a learned denoising network and can combine it with explicit application-specific priors.The framework introduces multiple variables for the network and additional regularization terms.
  • Algorithmic framework and prior stacking: Prior stacking in PDHG uses separate variables to incorporate the network and additional regularization J within one optimization scheme.J may itself consist of multiple priors, such as total variation.
  • Algorithmic framework and prior stacking: The algorithm’s parameterization includes data fidelity and regularization parameters, while step-size rescaling can eliminate an arbitrary γ under the stated transformation.The rescaling maps β to β/γ and α to α/γ when γ = cτ.
  • Deep convolutional denoising network: The denoising component is an end-to-end trained DnCNN-S-like network with 17 convolution layers, 3×3 kernels, and ReLU nonlinearities.Separate networks are trained for different Gaussian noise levels σ.
  • Deep convolutional denoising network: The denoising networks are evaluated against NLM and BM3D using average PSNR over 11 standard test images and different Gaussian-noise standard deviations.The same test images are used in the deconvolution experiments.

5. Evaluation

A fixed Gaussian-denoising network is tested without retraining on demosaicking and deconvolution, using explicit priors and parameter searches where needed. It achieves strong results across tasks, including performance close to specialized methods, while transfer from denoising is not uniformly predictable.

  • Evaluation setup: The same fixed denoising network is used for deconvolution and Bayer demosaicking, despite being trained only for Gaussian-noise removal.The network is not specifically trained for either inverse problem.
  • Demosaicking: Demosaicking achieves very high average PSNR and is surpassed only by a CNN specifically trained for demosaicking.The method also outperforms BM3D-based FlexISP* by about 1 dB in PSNR.
  • Demosaicking: PSNR varies by about 1.1 dB across differently trained networks, yet average demosaicking PSNR remains above 36 dB over a wide range of training noise levels.The evaluation also optimizes data-fidelity, TV, and cross-channel parameters by exhaustive grid search.
  • Deconvolution: Deconvolution uses five experiments on 11 images with Gaussian, squared, and motion blur under differing noise conditions.The comparison includes separately optimized parameters for each experiment.
  • Deconvolution: Using one network across all deconvolution experiments yields performance on par with other methods and similar to specialized networks in experiments a–d.The generic approach outperforms the specialized networks in experiment e, which removes motion blur.
  • Deconvolution: The denoising-network PSNR advantage over BM3D does not fully transfer to deconvolution, leaving only a comparably small PSNR difference.The paper identifies the conditions for fully transferring denoising performance to inverse problems as an open question.
  • Deconvolution: Deconvolution runtime averages approximately 2.5 s versus approximately 4 s for FlexISP*, a relative improvement of 37.5%.The comparison attributes the runtime difference to neural-network efficiency.
  • Parameter robustness: Deconvolution PSNRs remain stable across networks trained at different noise levels, indicating robustness to the particular network used.Very small training noise can perform poorly, whereas sufficiently large noise supports good results over a broad range.

6. Conclusion

The paper studies denoising neural networks as proximal operators and reports theoretical algorithmic equivalences alongside strong empirical reconstruction performance. It argues that fixed denoisers can reduce problem-specific retraining and support learned natural-image priors when training data is unavailable.

  • Four algorithms using neural networks as proximal operators have the same potential fixed-points.
  • PDHG step size merely rescales data-fidelity and other regularization parameters.
  • A fixed DnCNN-S network combined with PDHG and prior stacking achieved state-of-the-art results on demosaicking and deblurring, including robustness tests.
  • The approach may ease problem-specific retraining and provide learned natural-image priors when training data is unavailable.

Supplementary Material

The supplied passage identifies the affiliations associated with the supplementary material.

  • The listed affiliations are Technical University of Munich and University of Siegen.

Abstract

The supplementary material provides additional theoretical and experimental detail for the paper’s two exemplary reconstruction problems.

  • The supplementary material includes a proof of Remark 3.1 and additional information about the numerical experiments.
  • It evaluates image demosaicking and deconvolution with detailed qualitative and quantitative results.
  • Reported materials include grid-search parameter values, reconstruction PSNR values, and images.

Proof of Remark 3.1

The proof establishes that replacing the regularization proximal operator with a continuous function yields equivalent fixed-point equations across four algorithmic schemes. It also relates the PDHG2 scheme to ADMM under specified updates and parameters.

  • Replacing the proximal operator by an arbitrary continuous function G makes the fixed-point equations of PG, ADMM, PDHG1, and PDHG2 equivalent.
  • The proof derives the ADMM fixed-point relation through the optimality condition and initializes auxiliary variables to recover a fixed-point.
  • For PDHG1, the proof substitutes the relation y = −A^T ∇H_f(Au) into the update equations to obtain the common fixed-point equation.
  • For PDHG2, the same fixed-point relation follows from y = −A^T ∇H_f(Au), with initialization at the fixed-point.
  • PDHG2 is equivalent to ADMM in the convex case under overrelaxation, reversed update order, and specified parameters, and this remains valid after neural-network replacement.

Evaluation

The approach is evaluated on noise-free demosaicking and five deconvolution experiments using a fixed denoising network trained at σ = 0.02. Results are reported through residual visualizations, channel-wise and image-wise PSNRs, and parameter searches.

  • Demosaicking: The demosaicking evaluation uses 18 Bayer-filtered images from the McMaster color image dataset.The fixed denoising network was trained on noise with standard deviation σ = 0.02.
  • Demosaicking: The demosaicking results include magnified residual regions to illustrate reconstruction error across differently structured image areas.The images were cropped to avoid boundary effects.
  • Demosaicking: Channel-wise PSNR values are reported for each of the 18 McMaster images, with superior green-channel reconstruction attributed to the RGGB filter pattern.The green channel’s dominance in the RGGB pattern is given as the explanation for this result.
  • Deconvolution: The deconvolution evaluation comprises five experiments corrupting 11 standard test images with different blur kernels and Gaussian noise levels.The experiments report PSNR values for FlexISP∗ and multiple versions of the proposed approach.
  • Deconvolution: Deconvolution results compare image-wise PSNRs for FlexISP∗ and proposed variants using denoising networks trained on different σ values.The application-independent approach uses a network trained on σ = 0.02.
  • Parameter selection: Optimal data-fidelity and total-variation regularization parameters are obtained by extensive grid search for networks trained at different noise standard deviations.The PDHG dual step size is set to γ = 1.0, while the primal step size is determined subject to τγ < c.
Loading 1704.03488v2…