Source-linked AI summary
SNIPS: Solving Noisy Inverse Problems Stochastically
Bahjat Kawar, Gregory Vaksman, Michael Elad
TL;DR
Noisy linear inverse problems can admit many plausible reconstructions, while MMSE estimators may average them into perceptually weak images. SNIPS combines annealed Langevin dynamics, Newton's method, an MMSE Gaussian denoiser, and SVD-based score derivation to sample the posterior. It produces diverse, high-quality, measurement-consistent samples across deblurring, super-resolution, and compressive sensing, with reported gains over RED in deblurring.
Problem
Severely degraded inverse problems are ill-posed, and MMSE recovery can average plausible solutions into images lacking fine details and perceptual quality.
Method
SNIPS derives a tractable noisy posterior score using SVD decoupling and samples it with annealed Langevin dynamics, Newton's method, and a pre-trained Gaussian MMSE denoiser.
Results
SNIPS produces diverse, sharp, high-quality samples consistent with measurements across deblurring, super-resolution, and compressive sensing, outperforming RED by more than 11% in PSNR and 58% in LPIPS.
Takeaways & Limitations
Posterior sampling exposes uncertainty by returning multiple valid solutions for the same noisy observation rather than a single conditional-mean reconstruction.
Takeaways & Limitations
SNIPS requires costly SVD computation, does not currently handle general content images, and needs many iterations for sampling.
Abstract
from arXiv · showhide
In this work we introduce a novel stochastic algorithm dubbed SNIPS, which draws samples from the posterior distribution of any linear inverse problem, where the observation is assumed to be contaminated by additive white Gaussian noise. Our solution incorporates ideas from Langevin dynamics and Newton's method, and exploits a pre-trained minimum mean squared error (MMSE) Gaussian denoiser. The proposed approach relies on an intricate derivation of the posterior score function that includes a singular value decomposition (SVD) of the degradation operator, in order to obtain a tractable iterative algorithm for the desired sampling. Due to its stochasticity, the algorithm can produce multiple high perceptual quality samples for the same noisy observation. We demonstrate the abilities of the proposed paradigm for image deblurring, super-resolution, and compressive sensing. We show that the samples produced are sharp, detailed and consistent with the given measurements, and their diversity exposes the inherent uncertainty in the inverse problem being solved.
1 Introduction
The paper targets noisy linear inverse problems where MMSE-oriented recovery can lose detail, proposing stochastic posterior sampling to produce diverse, perceptually strong solutions. SNIPS combines Langevin-based sampling with an SVD-based treatment of the degradation operator and demonstrates results on several ill-posed image-recovery tasks.
- Noisy linear inverse problems include denoising, inpainting, deblurring, super-resolution, and compressive sensing.
- MMSE recovery averages multiple plausible clean images, often producing washed-out reconstructions that lack fine details under severe degradation.
- Posterior sampling is pursued instead of conditional-mean estimation to support diverse, measurement-consistent images with higher perceptual quality, while noisy-measurement handling remains important.
- SNIPS uses a carefully derived posterior score and annealing construction to generalize Langevin-based sampling to noisy linear inverse problems.
- The method samples from the posterior so repeated runs on one observation yield different valid solutions, demonstrated for deblurring, super-resolution, and compressive sensing.
2 Background
The background develops Langevin and annealed Langevin dynamics as stochastic sampling methods and explains how MMSE denoisers can approximate the required score. Prior inverse-problem extensions generally rely on simplified measurement models, motivating a more general treatment.
- Langevin dynamics samples from a distribution by iteratively following the gradient of its log density while injecting Gaussian noise.
- The injected noise preserves stochastic sampling, and under mild conditions the randomly initialized process converges to a sample from the target distribution.
- Annealed Langevin dynamics starts at high synthetic noise and gradually reduces it, with noise-dependent step sizes that improve convergence.
- A Miyasawa relation identifies the denoiser output as a conditional MMSE estimate, enabling denoising networks to replace the otherwise difficult score function.
- Earlier posterior-sampling methods for inverse problems typically assume noiseless measurements or simplified degradation operators, limiting direct generalization.
3 The Proposed Approach: Deriving the Conditional Score Function
SNIPS derives a tractable conditional posterior score for noisy linear inverse problems by transforming the problem into the SVD domain and carefully coupling synthetic and measurement noise. The resulting score combines measurement information with a denoiser-based prior score across singular-value-dependent regimes.
- Problem setting: SNIPS models recovery as sampling x from p(x|y) when y = Hx + z contains additive Gaussian measurement noise.The formulation assumes M ≤ N for convenience, while the approach also works for M > N.
- Synthetic annealing: Sampling is performed through blurred posteriors p(˜x|y), with noise levels σ decreasing from high values toward zero.The target ˜x = x + n uses synthetic noise whose construction is tied to the measurement noise.
- SVD-domain derivation: The SVD H = UΣV^T makes the conditional-score derivation tractable by separating the measurement operator into singular directions.The transformed variables yT = U^T y and ˜xT = V^T ˜x preserve the relevant information and support sampling in the SVD domain before transforming back with V.
- Noise construction: Synthetic noise is carved progressively from the measurement noise so the residual noise becomes Gaussian, uncorrelated, known in variance, and independent of ˜xi.The construction changes according to whether σisj is above or below the measurement-noise level σ0; zero singular-value directions remain independent of the measurements.
- Conditional score: The conditional score partitions singular directions into zero, weaker-than-measurement, and stronger-than-measurement regimes with distinct measurement and prior-score contributions.For zero singular values the score is measurement-independent; stronger directions are measurement-dependent, while intermediate directions combine blurred prior and measurement terms.
- Conditional score: The final score uses a denoiser- or neural-network-estimated prior score, while the remaining terms follow from the known operator and one initial SVD.The score vector ∇˜x log p(˜x) can be estimated with a neural network or a pre-trained MMSE denoiser.
4 The Proposed Algorithm
The proposed algorithm accelerates Langevin sampling with singular-value-aware step sizes inspired by Newton's method. Combined with the derived conditional score and a neural score estimator, it yields a tractable sampler that produces near-clean posterior samples across noisy inverse-problem settings.
- Motivation: Standard Langevin dynamics can converge extremely slowly because transformed coordinates advance at different speeds determined by their singular values.A very small global step size would otherwise be required for good performance.
- Adaptive updates: SNIPS replaces the scalar step size with a vector αi and diagonal matrix Ai to adapt updates across coordinates.The update retains Langevin noise while allowing coordinate-specific scaling.
- Newton-inspired scaling: The step sizes are derived from a diagonal approximation to the negative inverse Hessian, borrowing Newton's method to speed convergence.The Newton-inspired scaling is combined with the conditional score estimated from Equation 10.
- Sampling algorithm: Using the adaptive update, conditional score, and neural score estimator, SNIPS obtains a tractable iterative sampler for p(˜xL|y).At the final noise level, the remaining noise is sufficiently small to treat the output as a sample from the ideal image manifold.
- Special cases: The method specializes to image synthesis without measurements and to denoising or noisy inpainting for particular choices of H.With σ0 = 0, it also reduces to close variants of earlier iterative methods for selected operators.
5 Experimental Results
SNIPS is evaluated on noisy image deblurring, super resolution, and compressive sensing, producing diverse reconstructions while maintaining measurement faithfulness. Experiments also expose a perceptual-quality versus PSNR trade-off between individual samples and their empirical mean.
- Experimental setup: SNIPS is evaluated using NCSNv2 denoisers on CelebA and LSUN test sets, with 8 runs producing 8 samples for each input.Separate models are trained for CelebA 64×64 images, LSUN bedrooms 128×128 images, and LSUN towers 128×128 images.
- Image deblurring: Deblurring with a uniform 5 × 5 blur kernel and σ0 = 0.1 yields visually pleasing, diverse samples.The noise standard deviation refers to pixel values in the range [0, 1].
- Super resolution: Super-resolution tests use block averaging with 2 × 2 or 4 × 4 blocks and additive white Gaussian noise, with CelebA results shown for 4 : 1 and 2 : 1 downscaling.Figure 4 uses σ0 = 0.1 for 4 : 1 downscaling, and Figure 5 uses σ0 = 0.1 for 2 : 1 downscaling.
- Compressive sensing: Compressive sensing uses random projection matrices with singular values of 1 at compression levels of 25%, 12.5%, and 6.25%.More aggressive compression produces more significant reconstruction variations.
- Quantitative and perceptual assessment: Around 2.4 dB higher PSNR is achieved by the empirical conditional mean than by individual posterior samples, although the mean is less visually appealing.This is compared with the theoretically expected 3dB difference between posterior samples and the MMSE estimator.
- Quantitative and perceptual assessment: More than 11% improvement in PSNR and more than 58% improvement in LPIPS are reported for SNIPS over RED on deblurring results.LPIPS is used as a perceptual quality metric.
- Faithfulness to measurements: Measurement faithfulness is supported by residuals whose standard deviation nearly matches σ0, satisfy |ρ| < 0.1, and pass normality tests in around 95% of samples.The residual is evaluated as y − Hx̂, with p-values greater than 0.05 accepted as Gaussian.
6 Conclusion and Future Work
SNIPS is presented as a stochastic sampler for general noisy linear inverse problems, combining annealed Langevin dynamics and Newton’s method with a pretrained Gaussian MMSE denoiser. Its scope is demonstrated across three inverse-problem tasks, while scalability, image-content coverage, and runtime remain limitations.
- Conclusion: SNIPS samples high-quality, varied images from the posterior of general noisy linear inverse problems while remaining valid with respect to the measurements.Its derivation uses annealed noise and an SVD of the degradation operator to decouple measurement dependencies.
- Conclusion: SNIPS is demonstrated on image deblurring, super resolution, and compressive sensing.These tasks are described as noisy and ill-posed.
- Future work: SVD decomposition requires considerable memory and computation, hindering SNIPS’s scalability.The paper identifies this as the first limitation requiring attention in future extensions.
- Future work: The current SNIPS version does not handle general content images because of properties of the denoiser used.The paper links this scope boundary to the denoiser’s properties.
- Future work: Langevin-based sampling requires many iterations; producing 8 CelebA super-resolution samples takes 2 minutes in the reported tests.The authors suggest exploring acceleration methods.
A Conditional Score Derivation Proofs
The derivation obtains an approximate conditional posterior score by splitting SVD-domain variables into zero, greater-than, and less-than singular-value cases, then combining the resulting expressions. The derivation ultimately incorporates a blurred prior score estimated through denoising.
- Conditional score cases: The proof separates the conditional-score derivation into the zero, greater-than, and less-than singular-value components before concatenating them.Each component is handled with Bayes-rule manipulations and Gaussian score calculations.
- Zero singular values: For zero-singular-value components, the likelihood derivative vanishes, leaving the corresponding prior-score component.The argument uses the independence of the gradual noise additions from the relevant transformed variables.
- Greater-than components: For greater-than components, the derivation combines Gaussian likelihood terms with conditional prior-score terms and approximates denoiser expectations under an additional-information assumption.The assumption is that the extra zero-component information does not significantly change the estimation of the greater-than component.
- Less-than components: For less-than components, differentiating the Gaussian likelihood contributes a term involving the transformed measurement residual and the singular-value matrix.The likelihood contribution is derived using the Gaussian gradient-log and the inner derivative multiplied by −ΣT.
- Aggregated score: The three component results are aggregated into an approximate conditional score function with zeros in entries corresponding to the greater-than components.The resulting expression is the score used for the subsequent algorithmic construction.
B Step Size Derivation
The method derives a diagonal approximation to the log-posterior Hessian and uses its negative inverse as a position-dependent step-size matrix. This conditioning is motivated by the need to improve the Langevin iterations, while uniform step sizes can diverge under the tested deblurring settings.
- Hessian approximation: The Hessian of the log posterior is approximated by a diagonal matrix in the SVD domain.The derivation aggregates the three singular-value cases into diagonal Hessian entries.
- Step-size comparison: Figure 7 compares Ai with uniform time-dependent step sizes αi = c · σ2 i under fixed 5 × 5 blur and σ0 = 0.1 settings.The uniform choices use c = 1e −3 and c = 1e −5.
- Step-size construction: The negative inverse of the diagonal Hessian defines Ai, which supplies the position-dependent step size for the iterative Langevin updates.The diagonal entries are non-zero, so the approximation can be inverted directly.
- Step-size comparison: The uniform step size diverges under the same image-deblurring hyperparameters, whereas a sufficiently large iteration count might potentially yield convergence.The authors do not demonstrate that alternative because it would require retraining NCSNv2 and slow the algorithm.
C Alternative Definition of the Noise
The alternative noise definition makes the synthetic annealed Langevin noise independent of the measurement noise, but this independence prevents an analytical conditional-score derivation in the non-invertible case. The paper therefore uses an SVD-based noise construction to make the required score tractable.
- Alternative noise: With independent synthetic and measurement noise, the transformed variable has covariance involving σ2 0I + σ2 i HHT.The construction begins from ˜xL+1 = x and adds Gaussian noise through the annealing sequence.
- Alternative noise: Multiplication by H cannot convert p(˜xi|y) into p(H˜xi − y|y) when H is non-invertible because it changes the tested vector’s statistics.This blocks the direct conditional-score route used by the alternative formulation.
- Alternative noise: The likelihood term contains the Gaussian vector z − Hni conditioned on ˜xi, but the derivation lacks an explicit p(ni|˜xi) term.Consequently, the analytical gradient-log of the likelihood cannot be obtained.
- Adopted construction: The adopted construction uses the SVD of H and defines the noise additions so that y − H˜xi is independent of ˜xi.The paper states that both the SVD step and this noise sequence appear unavoidable for tractability.
D Implementation Details
Experiments use a decreasing geometric noise-level sequence and NCSNv2-compatible hyperparameters, with the inverse problem specifying H, σ0, and y. SNIPS activates a denoiser at each iteration and produces samples on a single RTX3080 GPU.
- Experimental setup: The noise levels {σi}L i=1 form a decreasing geometric sequence consistent with the NCSNv2 hyperparameters.The inverse problem determines H, σ0, and y.
- Sampling procedure: SNIPS runs τL overall iterations, activating a denoiser during each iteration.The cited implementation passage describes this as the completion of the sampling algorithm.
- Runtime: On one Nvidia RTX3080 GPU with 10GB memory, producing 8 samples for 64 × 64 CelebA images took around 2 minutes.The same passage reports around 6 minutes for another experiment, but its sentence is truncated.
- Reproducibility: The implementation code is available in the snips_torch GitHub repository.The paper provides the repository URL directly.
E Comparison to RED
SNIPS is compared with RED for deblurring using the same NCSNv2 denoiser, with PSNR and LPIPS used to assess fidelity and perceptual quality. SNIPS and its sample mean outperform RED on both metrics, while individual SNIPS samples offer better visual quality than their mean at high noise.
- SNIPS and its mean outperform RED in both PSNR and LPIPS for image deblurring.The comparison uses a uniform 5 × 5 blur kernel, additive noise, and the same NCSNv2 denoiser for both methods.
- At σ0 = 0.1, SNIPS has superior visual quality to its sample average, at the expense of PSNR.LPIPS is included alongside PSNR to evaluate perceptual quality, and Figure 8 provides a visual comparison.
F Additional Results
Additional experiments examine SNIPS on deblurring, super-resolution, and compressive sensing across CelebA and LSUN images. The figures vary degradation strength, compression level, dataset, and image category to display the resulting samples and their averages.
- Additional Results: Additional figures present SNIPS results for image deblurring, super-resolution, and compressive sensing.The accompanying guidance recommends zooming in to inspect details in individual samples and their average.
- CelebA: CelebA super-resolution experiments use 2:1 and 4:1 downscaling with additive noise σ0 = 0.1.Figure 10 additionally labels each set as original, low-resolution, and SNIPS restoration.
- CelebA: CelebA compressive sensing experiments use 25% compression with additive noise σ0 = 0.1.
- LSUN bedrooms: LSUN bedroom experiments cover 2:1 and 4:1 super-resolution and 25% compressive sensing, using additive noise σ0 = 0.04.
- LSUN towers: LSUN tower experiments cover 2:1 and 4:1 super-resolution and 25% compressive sensing, using additive noise σ0 = 0.04.