Source-linked AI summary
Posterior Information Dynamics of Diffusion Models for Linear Inverse Problems
Xiangming Meng
TL;DR
The paper asks when linear measurements remove uncertainty during diffusion reverse denoising and how remaining information is distributed across signal directions. It introduces the smoothed likelihood force and derives exact information identities, then evaluates solvable models and a frozen FFHQ illustration, finding that measurement information has both a budget and geometry not captured by endpoint quality or operator spectrum alone.
Problem
The reverse trajectory progressively resolves uncertainty, but endpoint quality does not show what uncertainty a linear measurement removes or how remaining uncertainty is distributed across signal directions.
Method
The paper analyzes the smoothed likelihood force, the posterior–prior score difference at each noise level, using relative de Bruijn, time-reversal, Girsanov, and I-MMSE identities.
Results
Solvable models and a frozen FFHQ illustration show that conditioning removes measurement-explained class separation, reduces empirical explanation entropy, and produces different null-space trajectory statistics for masks sharing a spectrum.
Takeaways & Limitations
Measurement information has both a budget and geometry, so endpoint quality and operator spectrum alone do not capture its assimilation during reverse dynamics.
Takeaways & Limitations
The biased-crossover analysis is specific to the symmetric zero-field reference and does not establish monotonicity for the baseline-corrected cloning observable at nonzero field.
Abstract
from arXiv · showhide
Diffusion models are widely used as priors for linear inverse problems, yet endpoint quality does not reveal when measurement information enters reverse denoising or how it is allocated across signal directions. We study this process through the smoothed likelihood force, the difference between exact posterior and prior scores at each noise level. For a fixed measurement, its expected squared norm gives both posterior--prior relative-entropy dissipation and reverse-path relative-entropy growth. Averaging over measurements yields an information--minimum mean-square error (I-MMSE) identity linking information gain to denoising-error reduction. Under finite second moments, the force energy and its ratio to prior-score energy decay quadratically in the noising kernel's signal coefficient at high noise. Solvable models show that conditioning removes class separation already explained by the measurement, reduces a uniform index entropy over \(n\) empirical samples from \(\log n\) to \(H(I\mid r)\), and makes assimilation depend on operator--prior alignment even for identical singular values. Experiments in models with tractable posteriors evaluate these predictions. In a separate illustration with a frozen FFHQ model, masks sharing the same spectrum yield different prior-normalized null-space trajectory statistics.
1 Introduction
The paper asks when linear measurements remove uncertainty during reverse diffusion and how the remaining information is distributed across signal directions. It develops exact information identities and evaluates these dynamics in solvable and learned-model settings.
- Endpoint reconstruction metrics do not reveal when measurement information enters the conditional diffusion trajectory or how it is distributed across signal directions.
- The smoothed likelihood force is the difference between posterior and prior scores at each noisy state, accounting for uncertainty in the clean signal.
- The force energy yields posterior–prior entropy-dissipation and reverse-path relative-entropy identities, while measurement averaging gives an I-MMSE relation to denoising-error reduction.
- Conditioning removes class separation already explained by the measurement and reduces empirical explanation resolution from log n to H(I | r).
- The study evaluates exact quantities in tractable models and examines null-space trajectory statistics for fixed-spectrum masks using a frozen FFHQ diffusion model.
2 Background and Problem Setting
The paper formulates noisy linear inverse problems as posterior inference under a diffusion prior. It defines shared forward noising, posterior and prior reverse processes, and operator-based coordinates for analyzing measurement effects.
- The forward process uses variance-preserving Ornstein–Uhlenbeck noising, producing a Gaussian-corrupted state at each positive noise level.
- Conditioning changes the initial clean law but leaves the forward noising mechanism unchanged, allowing posterior and prior marginals to evolve under the same diffusion semigroup.
- The smoothed likelihood force is the relative score between noised posterior and prior marginals and equals the measurement contribution to the posterior score.
- Operator geometry: Row-space and null-space projectors provide diagnostic coordinates, although the exact force need not remain row-supported for non-isotropic priors.
- Empirical priors and index uncertainty: For empirical priors, conditioning changes uniform support-index uncertainty, while state-resolved information tracks how much conditional index uncertainty has been resolved at noise level t.
3 Related Work
Prior work studies diffusion-based posterior sampling, likelihood approximations, operator-aware guidance, and trajectory-level errors. This paper instead characterizes the exact information law of the posterior–prior score difference before approximation.
- Diffusion priors have been used in likelihood-guided samplers, operator-aware constructions, pseudoinverse guidance, and variational or plug-and-play inverse solvers.
- Related methods include coordinate-wise spectral switching, sequential Monte Carlo, intermediate posterior sampling, covariance estimation, and conditional mutual-information objectives.
- Trajectory-level studies analyze bias, stability, or likelihood-weighted control for approximate conditional samplers and controlled paths.
- The paper’s complementary focus is the exact path-space relation between prior and posterior reverse laws generated by the same forward semigroup.
4 Foundational Identities and Posterior Force Budgets
The paper derives exact information identities for the smoothed likelihood force, linking fixed-measurement posterior–prior distinguishability to force energy and measurement-averaged denoising information.
- Posterior force: Conditioning and forward noising use the same diffusion semigroup, making the smoothed likelihood force the relative score between posterior and prior marginals.The force averages the clean-space likelihood over uncertainty in X0 given Xt, rather than evaluating it at a point estimate.
- Fixed-measurement information budget: The force energy equals the forward-time dissipation rate of posterior–prior relative entropy for each fixed measurement.This is the relative de Bruijn identity under the stated regularity and boundary assumptions.
- Reverse-path separation: Accumulated force energy, together with terminal marginal divergence, determines exact posterior–prior reverse-path relative entropy.The identity applies to reverse processes initialized from the exact time-T prior and posterior marginals.
- High-noise scaling: Under finite second moments, force energy and its ratio to prior-score energy decay quadratically in the noising kernel’s signal coefficient at high noise.The normalized statement uses a finite per-coordinate second-moment condition and a high-noise window with ∆t ≥ ∆0 > 0.
- Measurement-averaged information: Measurement averaging yields an I-MMSE representation in which information acquired under reverse denoising occurs at the measurement-force energy rate.The measurement-average information decreases under forward noising and is reparameterized through a Gaussian-channel signal-to-noise ratio.
- Scope: The high-noise bounds are joint-law averages and do not imply fixed-measurement monotonicity, a universal crossover time, or fixed-r Fisher orthogonality.At fixed r, the exact force-energy interpretation remains valid, but its high-noise coefficient may depend on r.
5 Posterior Information Mechanisms in Solvable Models
Three solvable models isolate how measurements reshape information resolved during diffusion: they remove already-explained class structure, reduce empirical index uncertainty, and allocate assimilation according to operator–prior alignment.
- Overview: Three solvable models isolate conditional class speciation, empirical sample-index resolution, and measurement-force allocation across operator–prior directions.Each model targets one mechanism rather than approximating a common data distribution.
- 5.1 Residual Class Information in Conditional Gaussian Mixtures: Measurement conditioning replaces class separation with the residual component left unresolved by the measurement, which alone can generate new class commitment.The measurement contributes a static class-evidence field, while the residual separation governs dynamic speciation.
- 5.1 Residual Class Information in Conditional Gaussian Mixtures: When the residual separation is nonzero, κt increases along the reverse trajectory and the zero-field posterior landscape changes at κt = 1.For nonzero measurement field, the symmetric pitchfork becomes a biased crossover, and baseline-corrected cloning measures information acquired beyond the measurement.
- 5.2 Empirical Explanation Resolution: A uniform empirical prior's explanation budget falls from log n to the conditional entropy H(I | r) after observing the measurement.The remaining budget is determined by the entropy of measurement-compatible indices, with weak measurements retaining a budget near log n.
- 5.2 Empirical Explanation Resolution: The cumulative posterior-resolution information approaches H(I | r) at the clean endpoint and is distinct from measurement-assimilation information Mγ = I(R; Yγ).This separates measurement-supplied information from explanation resolution supplied by the noisy state.
- 5.3 Directional Information Allocation and Operator–Prior Alignment: Equal operator singular spectra do not determine information gain, posterior contraction, or reverse-path assimilation under an anisotropic prior.At equal row-space budget, aligning measurements with leading prior eigendirections maximizes assimilation at every noise level.
6 From Exact Information Dynamics to Approximate Samplers
The paper connects exact information identities to approximate sampler diagnostics, then uses Gaussian-moment surrogates and denoiser-based observables when exact posterior quantities are unavailable.
- Measurement Assimilation and Index Resolution: A chain-rule identity partitions noisy-state index information into measurement assimilation and resolution beyond the measurement.The two terms have corresponding rate partitions and complementary clean-channel limits.
- Measurement Assimilation and Index Resolution: At the clean endpoint, total index information is log n, with log n − H(I | r) supplied by measurement and H(I | r) resolved during denoising.This separates the measurement head start from subsequent trajectory resolution.
- Gaussian-Moment Analysis of Approximate Samplers: For learned priors, a Gaussian approximation replaces X0 | Xt = x using the predictive mean and covariance to construct a surrogate measurement force.The surrogate is exact in the linear Gaussian model when the predictive covariance is state-independent.
- Gaussian-Moment Analysis of Approximate Samplers: The exact surrogate force includes covariance-state-dependence terms beyond the predictive-mean derivative.Residual-gradient guidance keeps the mean derivative but substitutes measurement variance for predictive covariance, producing a one-dimensional scale difference of 1 + vt/σ2_y before normalization.
- Approximate Sampler Comparisons: Approximate samplers differ in Jacobian treatment, step scaling, stochasticity, and correction placement, so comparisons concern complete fixed configurations.For learned FFHQ trajectories, denoiser residuals and row/null between-chain dispersion replace unavailable exact information quantities.
7 Diagnostic Studies
Exact-model diagnostics confirm the predicted high-noise scaling, conditional information budgets, directional force behavior, and covariance calibration; a frozen FFHQ study finds spectrum-matched masks with different null-space trajectory statistics.
- High-noise force decay: Log–log fits produce high-noise force-energy exponents between 1.002 and 1.008, close to the predicted unit exponent.Prior-score energy remains near one, and the leading force-energy coefficient agrees with d−1 Tr Cov(E[X0 | R]) to relative error below 4 × 10−4.
- Class-information diagnostics: In the conditional Gaussian mixture, measurement-induced fields replace the zero-field instability with a biased crossover and suppress only observed components of class separation.Subtracting the high-noise cloning baseline isolates evidence added by the noisy state.
- Empirical explanation resolution: The empirical-posterior entropy begins at H(I | r), not log n, then vanishes as components separate; lower-rank operators preserve more index ambiguity.Denoising leaves the smallest compatible set and strongest class baseline, while lower-rank operators retain more ambiguity.
- Reverse-trajectory diagnostics: The fixed-reference row-space force ratio peaks mid-trajectory and remains below 6 × 10−3 before decaying as empirical trajectories identify posterior-supported atoms.The low-noise decay occurs because posterior and prior denoisers converge to the same atom.
- Clean-estimate dispersion: The total-covariance calibration exchanges posterior-mean dispersion and remaining conditional covariance across row and null coordinates, with median relative gap 0.35%.The 95th-percentile relative gap is 3.21%, supporting dispersion as a resolved component.
- Neural trajectory diagnostics: Within one frozen FFHQ checkpoint and fixed-spectrum mask library, null-space dispersion AUCs differ across mask families, while row-space contrasts are smaller.The experiment does not identify a causal geometry effect because mask geometry and conditional task difficulty vary together.
8 Conclusion
The paper frames measurement conditioning as an information budget with geometric allocation, formalized through the smoothed likelihood force and tested in exact and learned-prior settings.
- Conclusion: The smoothed likelihood force is the posterior–prior score difference whose conditional energy gives entropy dissipation and reverse-path relative entropy.Averaging over measurements yields an I-MMSE relation, while high-noise force energy becomes weak relative to prior-score energy.
- Conclusion: Exact models show that conditioning removes class separation already explained by measurements, reduces empirical index uncertainty to H(I | r), and depends on operator–prior alignment.The operator–prior effect persists even when singular spectra are identical.
- Conclusion: A frozen FFHQ illustration finds different null-space trajectory statistics among masks sharing a spectrum.The study is a separate learned-model illustration rather than an exact evaluation of the information quantities.
- Conclusion: The conclusion is that endpoint quality and operator spectrum alone do not capture the budget and geometry of measurement information.Extending the analysis beyond linear measurements and testing directional predictions across independently trained models remain next steps.
- Scope and assumptions: The proof framework relies on shared forward noising, regularity and boundary assumptions, finite-entropy path identities, and finite second moments in its respective results.These assumptions support the entropy, force-energy, and high-noise arguments developed across the exact theory.
A.11 Proof of Proposition 4.11
The appendix proves the information identities and high-noise limits by combining Gaussian-channel I-MMSE relations, conditional-expectation orthogonality, and explicit Gaussian calculations for operator and class structure.
- Information identities: Conditional independence enables a chain-rule decomposition of mutual information into measurement and noisy-channel contributions.The resulting identities apply both unconditionally and conditional on the measurement.
- I-MMSE proof: The I-MMSE derivative expresses information growth through reduction in Gaussian-channel minimum mean-square error.Conditional versions are applied to the posterior input law for each measurement value.
- I-MMSE proof: Orthogonality separates unconditional denoising error into conditional denoising error plus the squared difference between conditional and unconditional posterior means.This yields the information-gain relation used in the proposition.
- Gaussian calculations: Gaussian conditioning gives explicit covariance blocks, posterior means, and force expressions for anisotropic linear measurements.The high-noise behavior follows from covariance expansions and bounded inverse operators.
- Gaussian mixture calculations: Class-conditional Gaussian calculations express posterior class evidence through measurement fields and the noised shifted coordinate.Posterior class weights and Gaussian likelihood ratios combine to produce the class-evidence formula.
A.15 Proof of Proposition 5.4 and Corollary 5.5
The proof reduces the conditioned Gaussian-mixture analysis to a single informative coordinate and characterizes when posterior geometry changes from unimodal to double-well. It also shows that conditioning removes class information already explained by the measurement.
- Conditional class information: P(C = c | Z_t, r) converges to P(C = c | r), so high-noise states retain only class information unexplained by the measurement.The conditional limiting class probabilities equal the measurement-conditioned class weights α_c.
- Conditional class information: Conditionally, the class posterior depends only on the whitened coordinate aligned with the residual class separation; orthogonal coordinates carry no class information.The informative coordinate has Gaussian means ±√κ_t and unit variance.
- Posterior transition: The posterior Hessian at the origin is negative definite for κ_t < 1, singular at κ_t = 1, and has a positive direction for κ_t > 1.The boundary case is resolved using the quartic term in the local expansion.
- Posterior transition: For κ_t > 1, the origin becomes a local minimum in one dimension and a saddle in higher dimensions, while orthogonal directions retain negative curvature.The transition corresponds to formation of a double-well profile along the informative direction.
- Reverse-time evolution: κ_t is nondecreasing along the reverse trajectory and strictly increases when the residual class separation is nonzero; if δ = 0, no instability occurs.The threshold κ_t = 1 is crossed only when the measurement leaves class separation unresolved.
A.18 Proof of Theorem 5.8
The theorem proof analyzes posterior and prior score energies coordinatewise in a common eigenbasis and establishes monotone reverse-time assimilation. It then connects Gaussian information gain and covariance contraction to operator–prior alignment.
- Directional decomposition: In the common eigenbasis, each signal coordinate is treated as a scalar Gaussian channel whose posterior and noised variances determine the score-energy ratios.The noised prior variance is B_i,t, while conditioning produces the corresponding posterior variance.
- Reverse-time monotonicity: The directional ratio ρ_i is strictly monotone in reverse time, and ρ_i = 1 is exactly the variance-halving condition.An interior crossing is unique when the relevant singular value is nonzero.
- Operator–prior alignment: Gaussian mutual information and posterior covariance-volume ordering are monotone in the relevant spectral quantities, with equality characterized by invariant extremal eigenspaces.These bounds use determinant inequalities for matrix compressions.
- Row/null decomposition: Global and row-space ratios share the same numerator when the force has no null-space component, so their difference is controlled by their score-energy denominators.This equality relies on the stated row-support hypothesis.
- Denoising decomposition: The total-covariance identity splits posterior covariance into expected within-state covariance and covariance of conditional clean-state estimates.The decomposition is obtained by conditioning on the noisy state X_t.
- Denoising decomposition: At zero channel SNR, the between-state term vanishes; at infinite SNR, the within-state term vanishes and the projected posterior covariance remains.The conditional denoiser converges in L2 to E[X_0 | r] at zero SNR.
B.1 Exact-Model Experiments
The exact-model experiments evaluate entropy, class, covariance, and high-noise scaling diagnostics using tractable MNIST posteriors. They report accurate force-energy asymptotics, stable reverse discretization, and likelihood-width sensitivity.
- Discretization checks: The 200/400/800-step comparison shows stable half-times and residual floors under reverse Euler–Maruyama discretization.The index half-time varies by at most 0.059 within a task, and the terminal residual stays within 6.3% of the analytic floor.
- High-noise scaling: Force-energy slopes are 1.003, 1.006, and 1.008, while relative-energy slopes are 1.002, 1.005, and 1.007 across denoising, super-resolution, and inpainting.The prior-score energy per dimension remains between 0.996 and 1.014.
- High-noise scaling: The estimated leading coefficient agrees with d^-1 Tr Cov(E[X_0 | R]) to relative error below 4 × 10^-4 in every task.This directly tests the predicted high-noise coefficient.
- Likelihood-width sensitivity: As likelihood width increases from 0.5 to 3, posterior index budgets approach log n and class baselines approach their unconditional value.The main MNIST diagnostics use σ_y = 2.
- Timing calibration: The Gaussian resolution scale has mean absolute error 0.109 versus 0.133 for a cross-validated constant predictor, but its advantage remains unresolved.The paired difference is −0.024 with a reference-cluster 95% interval from −0.064 to 0.008.
B.2 Neural-Sampler Experiments
The neural-sampler experiments compare common trajectory diagnostics across diffusion inverse-problem methods and fixed-spectrum masks. They find mask-dependent null-space behavior, while finite-chain checks support aggregate but not reference-level peak comparisons.
- Evaluation design: The dispersion summaries apply identical clipping and explicit row/null projectors across methods, while the full sampling rules differ in several simultaneous components.The configuration comparison therefore concerns complete samplers rather than one isolated implementation choice.
- Fixed-spectrum masks: All masks hide one quarter of spatial positions and therefore share the same rank, nullity, and singular values.The library includes contiguous, multi-hole, structured-distributed, and random masks.
- Mask-dependent trajectories: Contiguous-minus-random null-space contrasts have the same positive sign for all four method-specific comparisons.Multi-hole masks are intermediate, while structured-distributed masks are close to the random family.
- Mask-dependent trajectories: Endpoint missing-region PSNR is lower for contiguous masks than for random masks for every sampler.The comparison is reported alongside the null-space trajectory contrasts.
- Finite-chain precision: Relative to 64 chains, eight chains have pooled mean absolute relative errors of 8.0% for row-space AUC and 7.1% for null-space AUC.Peak errors are larger, so aggregate AUC comparisons are supported while reference-level peak rankings remain imprecise.
- Guidance calibration: The plug-in gradient uses measurement variance instead of predictive variance, and this discrepancy persists even with an exact clean denoiser.Residual normalization reduces the fixed-precision mismatch but does not reproduce exact-force predictive-variance damping.