Source-linked AI summary
Bayesian imaging using Plug & Play priors: when Langevin meets Tweedie
Rémi Laumont, Valentin de Bortoli, Andrés Almansa, Julie Delon, Alain Durmus, Marcelo Pereyra
TL;DR
Plug & Play Bayesian imaging lacks fully established links between denoisers, well-defined Bayesian models, and convergent sampling or optimization schemes. This paper develops a formal M-complete framework and convergence-guaranteed algorithms, demonstrating point estimation and uncertainty analysis on image restoration tasks while exposing limitations from denoiser regularization and chain convergence.
Problem
Plug & Play methods lack clear guarantees that denoiser-defined Bayesian models and sampling schemes are well defined, well posed, and convergent.
Method
The paper analyzes Plug & Play priors through M-complete Bayesian modeling and develops two Plug & Play ULA methods with convergence guarantees under realistic denoiser assumptions.
Results
The framework supports point estimation and uncertainty visualization and quantification in deblurring and inpainting experiments, with theoretical guarantees for the proposed algorithms.
Takeaways & Limitations
Plug & Play Bayesian inference can be given a well-defined theoretical framework while retaining denoiser-based priors and enabling uncertainty analysis.
Takeaways & Limitations
Slow Markov-chain convergence and poor regularization of some frequencies can slightly deteriorate MMSE solutions and produce rotated rectangular artifacts.
Abstract
from arXiv · showhide
Since the seminal work of Venkatakrishnan et al. in 2013, Plug & Play (PnP) methods have become ubiquitous in Bayesian imaging. These methods derive Minimum Mean Square Error (MMSE) or Maximum A Posteriori (MAP) estimators for inverse problems in imaging by combining an explicit likelihood function with a prior that is implicitly defined by an image denoising algorithm. The PnP algorithms proposed in the literature mainly differ in the iterative schemes they use for optimisation or for sampling. In the case of optimisation schemes, some recent works guarantee the convergence to a fixed point, albeit not necessarily a MAP estimate. In the case of sampling schemes, to the best of our knowledge, there is no known proof of convergence. There also remain important open questions regarding whether the underlying Bayesian models and estimators are well defined, well-posed, and have the basic regularity properties required to support these numerical schemes. To address these limitations, this paper develops theory, methods, and provably convergent algorithms for performing Bayesian inference with PnP priors. We introduce two algorithms: 1) PnP-ULA (Unadjusted Langevin Algorithm) for Monte Carlo sampling and MMSE inference; and 2) PnP-SGD (Stochastic Gradient Descent) for MAP inference. Using recent results on the quantitative convergence of Markov chains, we establish detailed convergence guarantees for these two algorithms under realistic assumptions on the denoising operators used, with special attention to denoisers based on deep neural networks. We also show that these algorithms approximately target a decision-theoretically optimal Bayesian model that is well-posed. The proposed algorithms are demonstrated on several canonical problems such as image deblurring, inpainting, and denoising, where they are used for point estimation as well as for uncertainty visualisation and quantification.
1 Introduction
Imaging inverse problems are often ill-posed, motivating Bayesian priors and Plug & Play methods that combine explicit likelihoods with denoiser-defined priors. This paper develops a formal Bayesian framework and convergence-guaranteed Plug & Play algorithms for estimation and sampling.
- Algorithmic foundations: The proposed Plug & Play ULA methods are related to Moreau-Yosida ULA but replace its proximal construction with a state-of-the-art Gaussian denoiser.Tweedie’s identity is used to relate the prior gradient to an MMSE denoiser.
- Plug & Play Bayesian imaging: Plug & Play methods combine an explicit likelihood with a prior implicitly represented by a denoising algorithm for MAP or MMSE inference.They can use denoiser-derived approximations of either the prior score or proximal operator within optimization or Monte Carlo schemes.
- Open challenges: Because denoisers are generally not directly linked to marginal distributions, the resulting Bayesian models and sampling schemes may lack clear definitions and convergence guarantees.The paper identifies this as a central challenge in plugging denoisers into gradient-based algorithms such as ULA.
- Paper contributions: The paper presents a formal framework for analyzing Plug & Play priors, including whether their models and estimators are well defined, well posed, and sufficiently regular for computation.The framework uses M-complete Bayesian modeling to interpret Plug & Play models relative to an unknown optimal marginal and posterior distribution.
- Theory and experiments: The paper establishes detailed convergence guarantees for two Plug & Play ULAs under realistic denoiser assumptions, including for a denoiser based on a deep neural network.Experiments cover non-blind image deblurring and inpainting, with point estimation, uncertainty visualization, and comparisons with Plug & Play SGD.
2 Bayesian inference with Plug & Play priors: theory methods and algorithms
The paper formalizes Plug & Play Bayesian inference by smoothing an unknown oracle prior, relating denoisers to posterior gradients through Tweedie’s identity, and establishing well-posedness and convergence conditions. The resulting regularized models can approximate the oracle arbitrarily closely while supporting gradient-based computation under explicit assumptions.
- Bayesian modelling: PnP Bayesian models are analyzed as operational approximations to an unknown, decision-theoretically optimal oracle prior and posterior.The M-complete viewpoint treats practitioner-specified models as approximations rather than true models.
- Bayesian modelling: Gaussian smoothing constructs a proper, smooth regularized prior and posterior whose approximation error vanishes as ε →0.The regularized posterior is obtained by combining the smoothed prior with the likelihood under stated regularity assumptions.
- Algorithms and guarantees: When Gaussian denoising has uniformly bounded expected MSE, the regularized prior score is Lipschitz, enabling convergent Langevin-based Bayesian computation.The paper notes that weaker finite-MSE conditions give only local Lipschitz continuity and are left for future technical development.
- Algorithms and guarantees: The resulting PnP-ULA Markov chain has an invariant distribution whose density is provably close to the regularized oracle target under step-size conditions.This provides the theoretical basis for approximate sampling using the denoiser-based recursion.
- Well-posedness: Under mild likelihood assumptions, the regularized posterior is well posed, locally Lipschitz in the observation, and yields a stable MMSE estimator.The paper establishes local Lipschitz continuity with respect to y under an appropriate probability metric.
- Denoiser-based inference: Tweedie’s identity explicitly links the regularized posterior gradient to a denoising operator, making denoiser accuracy central to PnP inference accuracy.The framework compares an implemented denoiser with the oracle MMSE denoiser and characterizes approximation quality through their closeness.
3 Theoretical analysis
The theoretical analysis establishes when PnP-ULA and its projected variant are stable, geometrically convergent, and close to the intended smoothed Bayesian targets under explicit assumptions on likelihoods, denoisers, step sizes, and constraint sets.
- 3.2 Convergence of PnP-ULA: PnP-ULA is analyzed through geometric ergodicity and quantitative total-variation bounds between its iterates, invariant distribution, and the target πε.The analysis first establishes Markov-chain convergence and then controls the discrepancy between the stationary distribution and πε.
- 3.2 Convergence of PnP-ULA: Under H4, the PnP-ULA kernel admits an invariant probability measure πε,δ, enabling bounds on finite-sample Monte Carlo bias under moment and regularity conditions.The resulting bounds apply for convex compact constraint sets containing an appropriate ball and for step sizes below the stated threshold.
- 3.2 Convergence of PnP-ULA: Lipschitz control of the denoiser supports stability and geometric convergence, and can be enforced during neural-network training through weight regularization.The approximation error between Dε and the oracle denoiser is also part of the required conditions.
- 3.2 Convergence of PnP-ULA: Under H1, H2(R), and H3, the Markov chain is geometrically ergodic in W1 and a V-norm with V(x)=1+∥x∥2, using sufficiently small λ and δ.The contraction constants depend on model and constraint parameters but not directly on dimension.
- 3.2 Convergence of PnP-ULA: When the likelihood is not strongly concave, a user-fixed convex compact set C stabilizes the chain, while projection is theoretically important even though experiments rarely activate it.For noninvertible A, including deblurring kernels with Fourier zeros, strong concavity may fail; in experiments, PnP-ULA and PPnP-ULA therefore behave identically when projection is inactive.
4 Experimental study
Experiments evaluate PnP-ULA and PPnP-ULA on deblurring and inpainting, examining convergence, point-estimation quality, and uncertainty. Sampling quality varies substantially across images, with multimodality and anisotropy producing slow mixing and higher computational cost.
- Inpainting: Inpainting uncertainty is highest near edges, textured regions, and complex structures, while homogeneous regions have low marginal posterior standard deviation.These maps expose where the posterior remains uncertain in the reconstructed image.
- Inpainting: Inpainting chains generally mix through fluctuations around the MMSE, but Simpson exhibits metastability with transitions between regions after millions of iterations.Slow ACF decay indicates slow movement and high-variance Monte Carlo estimates.
- Quantitative results: Over 10^6 iterations are required for slowly converging cases, compared with approximately 10^5 for fast-converging experiments; larger-step PPnP-ULA can reduce the iterations needed for stable posterior means.The slow cases include deblurring of Alley, Bridge, and Goldhill.
- Point estimation: MMSE estimates from PnP-ULA are sharper and achieve better PSNR/SSIM than the best MAP results for the first three deblurring images.For the remaining images, slow mixing and poor frequency regularisation deteriorate MMSE quality and produce a rotated rectangular artefact.
- Point estimation: Sampling can outperform MAP estimation when it correctly explores the posterior, but this quality improvement requires substantially more computation.The experiments also use posterior samples to visualise uncertainty across pixel and larger spatial scales.
5 Conclusion
The paper establishes a mathematically grounded framework for Bayesian inference with Plug & Play priors, with convergence guarantees and experiments demonstrating point estimation and uncertainty analysis.
- The framework establishes well-defined, well-posed Plug & Play Bayesian models and convergence guarantees for three algorithms related to biased Langevin diffusion.The analysis does not require denoisers to be gradient or proximal operators and also controls estimation error relative to an intractable oracle model.
- Deblurring and inpainting experiments compute point estimates, visualize and quantify uncertainty, and expose how denoiser limitations affect the resulting Bayesian model.
- Future work includes analytic-plus-denoiser priors, generative and autoencoder priors, alternative smoothings, accelerated algorithms, and medical-imaging uncertainty quantification.
- The supplementary material extends the framework, strengthens convergence results for strongly log-concave likelihoods, and provides posterior approximation bounds and proofs.
C Strongly log-concave case
In the strongly log-concave setting, the analysis proves geometric ergodicity and finite-sample bounds for the Markov chain and its Monte Carlo estimators under explicit step-size and regularity conditions.
- The strong-concavity condition applies to denoising with A=Id and to deblurring kernels having full Fourier support.
- Strong log-concavity yields geometric ergodicity of the Markov chain in Wasserstein and weighted norms under the stated assumptions.
- For sufficiently small step sizes, the chain admits an invariant probability measure and satisfies an explicit convergence bound with a Lyapunov function V(x)=1+∥x∥2.
- The resulting Monte Carlo estimator has a finite-sample bias bound for measurable test functions with controlled linear growth.
D Posterior approximation
This section establishes conditions under which the smoothed Plug & Play posterior approximates the target posterior and derives finite-sample bounds under general and strongly log-concave settings.
- Hölder regularity and coercive growth conditions on the prior potential provide easy-to-check sufficient conditions for the required translation regularity assumption H6(α).
- For α⩾1 and bounded prior density, the smoothed posterior πε converges to the target posterior π in total variation as ε→0.
- Under H1, H5, H2, H4, and H3, the general posterior approximation result combines smoothing, discretization, truncation, and finite-iteration error terms.
- In the strongly log-concave case, the bound simplifies under an explicit concavity condition and contains the terms ε^β min(α,1)/4, δ^1/2, M_R, exp[-R], and (nδ)^-1.
E Technical results
The technical results establish drift conditions and convergence properties for Langevin discretizations and their continuous-time limits, supporting quantitative control of the associated Markov processes.
- The Langevin recursion defines a Gaussian-noise Markov kernel whose convergence analysis is based on discrete drift conditions and continuous-time semigroup properties.
- Under strong convexity and Lipschitz-gradient conditions, the discretized kernel and diffusion satisfy discrete and continuous drift conditions for polynomial Lyapunov functions.
- The technical proof constructs contraction and drift bounds through inequalities controlling the update map, moments, and exponential Lyapunov functions.
- The resulting estimates support quantitative comparisons between discretized chains and diffusions, including bounds involving iteration count and step size.
- Additional results establish local Lipschitz dependence of the posterior on the observation in total variation under integrability conditions.
F Proofs of Section 3.2
This section establishes convergence and invariant-measure bias control for PnP-ULA under two posterior regularity regimes.
- The Markov chain uses a stepsize, algorithmic hyperparameters, projection onto a closed convex set, and i.i.d. standard Gaussian noise.
- The analysis treats PnP-ULA in a general framework with α ≠ 1 and controls the bias of its invariant measure.The results cover log-concave posteriors and posteriors satisfying a one-sided Lipschitz condition; the latter statements are given for α = 1.
F.1 Proof of Proposition 4
This proof develops the local regularity argument needed for Proposition 4, using smoothed densities and Lipschitz properties of the denoising-related quantities.
- The proof introduces Gaussian-smoothed variables and their densities, including the conditional density of X given its noisy observation.
- Local Lipschitz control is established on bounded balls for the relevant denoising-related map.
- Volume estimates, Fubini’s theorem, Lebesgue differentiation, and dominated convergence complete the bounded-region argument.
F.2 Proof of Proposition 5 and Proposition 10
These proofs establish contraction-related bounds for PnP-ULA under one-sided Lipschitz and log-concavity assumptions on the posterior.
- Under a one-sided Lipschitz condition, the drift becomes strongly dissipative outside a radius determined by the constraint-set diameter.For distances at least 4R_C, the drift inner product is bounded above by -∥x_1 − x_2∥^2/(4λ).
- The proof combines these drift bounds with quantitative Markov-chain results to conclude Proposition 5.
- For an m-concave log-posterior satisfying 2αL/(mε) ⩽ 1, the corresponding contraction argument concludes Proposition 10.
F.3 Proof of Proposition 6 and Proposition 11
These proofs establish regularity of the smoothed posterior, compare diffusion processes through Girsanov’s theorem, and derive invariant-measure results for PnP-ULA.
- The proof compares diffusion laws using Girsanov’s theorem and generalized Pinsker inequalities under stated integrability assumptions.
- Under H4, the gradient of the smoothed log-density is Lipschitz continuous, and the smoothed log-density is infinitely differentiable.
- The remaining estimates control discretization and approximation terms using moment bounds, semigroup arguments, and Markov-chain lemmas.
- The constructed diffusion has a Lipschitz drift, admits a unique strong solution, and defines a Markov semigroup with quantitative contraction properties.
- A fixed-point argument yields invariant probability measures for both the discrete chain and the continuous-time diffusion.
G Proofs of Section 3.3
The proof establishes a contraction-based convergence bound for a coupled Markov chain by combining one-step coupling estimates with the absorbing property of the diagonal.
- Coupling estimate: The coupling construction yields a bound on the probability that the two Markov-chain components have not coalesced after ℓ steps.The chain starts from two points in the compact convex set C, and the argument invokes a coupling result from [16].
- Absorption: The diagonal Δ_Rd is absorbing once the coupled components meet, allowing the coalescence estimate to extend to arbitrary k.The proof explicitly uses the absorbing property to relate later-time behavior to the coupling event.
- Geometric bound: Combining the coupling estimate with the absorbing property produces a geometric bound for every k ∈ N.The intermediate bound is combined with equation (37) before the final rate is identified.
- Rate: The final contraction parameter is expressed as ˜ρ_C = (1 − β)^(1/(1+¯δ)), after converting the integer block count using a floor-function inequality.The resulting rate depends on β and ¯δ and is obtained by the estimate involving ⌊k/⌈1/δ⌉⌋.
G.2 Proof of Proposition 9
The proof of Proposition 9 combines drift, moment, and maximal-inequality arguments to control the Markov-chain Lyapunov function uniformly over the relevant parameters.
- Assumptions: The step-size assumptions constrain λ, α, and ε through 2λ(L_y + αL/ε − min(m, 0)) ≤ 1 and define an admissible upper bound ¯δ_1.These conditions are imposed before applying the subsequent drift estimates.
- Drift control: A coercive inner-product bound and a Markov-chain theorem provide constants controlling the Lyapunov function over iterations.The resulting constants ˜B and ˜ρ are chosen independently of R.
- Uniformity: The proof obtains a uniform recursive bound with constants ˜λ and ˜c that do not depend on R.This bound is stated for δ ≤ ¯δ_2 and is used in the later maximal-probability argument.
- Supermartingale argument: Choosing A = λ exp[cλ − ¯δ_2] yields R_ε,δV(x) ≤ AδV(x), so the normalized Lyapunov sequence is a supermartingale.Doob’s maximal inequality and Markov’s inequality are then applied to the supermartingale.
H.1 Proof of Proposition 13
The proof of Proposition 13 develops integrability and approximation bounds using coercivity, Gaussian mollification, and elementary power inequalities.
- Integrability: A lower bound U(x) ≥ c_1∥x∥^ϖ + c_2 supplies coercive growth used to control integrability for the relevant powers and parameters.The argument applies this condition for k ∈ N* and α > 0.
- Auxiliary inequalities: The proof introduces a technical lemma and establishes the power-difference inequality needed for later estimates.For β > 0, the lemma bounds (x + y)^β − x^β by terms involving y^β and x^((β−1)∧0)y.
- Measure comparison: Lemma 22 provides a comparison framework for probability measures whose densities are normalized versions of measurable nonnegative functions.This lemma is invoked in the proof of Proposition 14 before the approximation estimates are combined.
- Mollification: Gaussian mollification approximates p in L1 as ε tends to zero, which is combined with equation (39) and Lemma 22 to conclude the first part.The proof uses that Gaussian kernels form a family of mollifiers and explicitly obtains lim_{ε→0} ∫|p(x) − p_ε|dx = 0.
- Remaining estimates: Jensen’s inequality, a Gaussian rescaling, and boundedness of p_ε support the remaining estimates under the assumption H6(α).The argument treats α ≥ 1 using equation (38), while the bound ∥p_ε∥∞ ≤ ∥p∥∞ is used for ε > 0.