Source-linked AI summary

Physics-Driven Independent Pair Generation for Iterative Self-Supervised Low-Dose CT Denoising

Xianlei Han, Shaoyu Wang, Jiancheng Fang, Weiwen Wu, Qiegen Liu

arXiv:2609.02654v1cs.CV

TL;DR

LDCT self-supervised denoising lacks explicit treatment of mixed Poisson–Gaussian noise and independent pair requirements. The proposed physics-driven framework infers photon counts, constructs variance-matched pairs, and iteratively refines a sinogram prior with image-domain feedback. Across simulated and real LDCT data, it improves over evaluated self-supervised baselines and approaches the evaluated supervised baseline, subject to model and estimation assumptions.

  • Problem

    Self-supervised LDCT methods often rely on generic image statistics without explicitly modeling mixed Poisson–Gaussian noise, while paired normal-dose data are difficult to acquire clinically.

  • Method

    A learned sinogram prior and the CT noise model separate photon and Gaussian components for thinning-based pair construction, residual variance matching, and cross-domain iterative refinement.

  • Results

    The proposed method achieves the best PSNR, SSIM, and RMSE among self-supervised methods on all datasets at both dose levels and outperforms supervised RED-CNN in 17 of 18 entries.

  • Takeaways & Limitations

    Physics-driven pair construction combined with cross-domain iteration provides a self-supervised LDCT denoising framework that is effective across dose levels and datasets.

  • Takeaways & Limitations

    Practical branch independence may be compromised by posterior and prior-estimation errors or substantial deviations from the assumed Poisson–Gaussian acquisition model.

Abstract

from arXiv · show

Low-dose computed tomography (LDCT) measurements contain mixed Poisson-Gaussian noise. However, most self-supervised methods rely on generic image statistics and do not explicitly model this noise, which may limit their ability to effectively suppress realistic LDCT noise. To address this issue, we propose a physics-driven framework with cross-domain iteration for self-supervised LDCT denoising. The proposed framework proceeds in three main steps. First, a learned sinogram prior and the LDCT noise model guide posterior inference of photon counts, enabling separation of the Poisson and Gaussian components. Second, the separated Poisson and Gaussian components are respectively processed by binomial thinning and Gaussian data thinning to construct two branches, and residual scaling matches each branch's noise level to that of the observation, yielding a training pair with approximately independent noise realizations from one low-dose measurement. Finally, the pair is used to train an image-domain network whose forward-projected outputs update the prior. Through cross-domain iteration, the prior and the training pair are progressively refined while maintaining consistency with CT acquisition physics. Experiments on simulated data from AAPM, LIDC-IDRI, and LoDoPaB-CT and on real LDCT data show consistent gains over the evaluated self-supervised baselines across dose levels, with performance comparable to the evaluated supervised baseline.

I. INTRODUCTION

LDCT self-supervised denoising must account for mixed Poisson–Gaussian acquisition noise while producing independent training pairs. The proposed framework addresses this through physics-driven pair construction and cross-domain iteration.

  • Motivation: LDCT combines increased mixed Poisson–Gaussian noise with clinically difficult paired normal-dose data requirements.These factors motivate self-supervised alternatives to supervised denoising.
  • Limitations of Existing Methods: Existing self-supervised methods operate mainly in image or projection domains but rarely jointly model mixed noise, acquisition physics, and cross-domain information.Projection partitioning can also alter effective dose and mismatch training and inference noise statistics.
  • Proposed Framework: The framework infers latent photon counts with a learned sinogram prior, separates noise components, and constructs approximately independent pairs from one low-dose observation.Binomial and Gaussian data thinning process the separated components into two stochastic branches.
  • Cross-Domain Iteration: Cross-domain iteration uses image-domain feedback to refine the sinogram prior, improving pair reliability and physical consistency between projection and image domains.This changes pair construction from a static procedure into a progressively updated modeling process.
  • Evidence: Theoretical analysis establishes pair independence and variance matching under ideal Poisson–Gaussian conditions, while experiments validate effectiveness across dose levels and datasets.The experiments cover simulated and real LDCT settings described in the paper context.

III. METHODOLOGY

The methodology models LDCT projection measurements as mixed Poisson–Gaussian observations, learns a sinogram prior, and constructs variance-matched branches for self-supervised training. Masking strategies and cross-domain feedback refine the prior and preserve acquisition consistency.

  • Noise Modeling: Each LDCT projection measurement is modeled as a Poisson photon count plus independent Gaussian readout noise.Y is the low-dose projection measurement, C ∼ Poisson(µ), and ε ∼ N(0, σ2).
  • Noise Modeling: Photon-count fluctuations and electronic readout noise are handled separately to construct projection-domain pairs consistent with CT acquisition physics.Forward-projected restored images update the prior because prior errors can propagate into generated samples.
  • Sinogram-Domain Prior Learning: The sinogram prior predicts masked pixels using contextual information, with a mapping operator restoring blind-spot predictions to their original coordinates.Each sinogram pixel is predicted once per forward pass.
  • Sinogram-Domain Prior Learning: A non-blind branch and re-visibility objective compensate for information excluded by blind-spot masking while constraining blind-spot error propagation.The visible-branch contribution increases during training, transitioning toward greater use of visible information while preserving details.
  • Sinogram-Domain Prior Learning: High and low masking ratios balance robust masked prediction against richer context and detail recovery during training.The schedule shifts emphasis from the stronger high-ratio blind-spot constraint to the richer context of the low-ratio branch.

C. Posterior-Guided Pair Construction

Posterior-guided pair construction separates latent Poisson counts from Gaussian readout noise, then generates and variance-matches two approximately independent branches for image-domain training.

  • Posterior inference: A learned sinogram prior and the mixed-noise model support posterior inference of latent photon counts from low-dose measurements.The posterior combines a Poisson count prior with a Gaussian readout-noise likelihood.
  • Posterior correction: Posterior mean and variance summarize the inferred count and its uncertainty, while discrepancy-based correction limits unreliable prior estimates.The adaptive correction moves the count prior toward the posterior estimate when disagreement is larger.
  • Branch construction: Poisson and Gaussian components are separately thinned to produce two stochastic branches from one low-dose observation.Binomial thinning handles counts, while independently sampled Gaussian noise produces the Gaussian branches.
  • Variance matching: Residual scaling compensates for logarithmic-transform variance inflation so branch noise matches the observation level used during inference.The variance-matching coefficient is k = √d_k, changing residual amplitude while preserving stochastic branch generation.
  • Reconstruction: The scaled sinograms are reconstructed into two images that form the image-domain self-supervised training pair.The resulting branches preserve separate random components under the stated ideal assumptions.

D. Cross-Domain Iterative Training

Cross-domain iterative training alternates pair generation, image restoration, and forward projection so restored images update the sinogram prior and promote consistency between domains.

  • D. Cross-Domain Iterative Training: The corrected sinogram prior initializes iterative training and conditions generation of two stochastic image branches.The generator uses posterior-sampling and pair-construction randomness to produce the branches.
  • D. Cross-Domain Iterative Training: The image-domain network is optimized on reconstructed branch images, after which restored estimates are averaged and forward-projected.The forward projector A(·) produces the reference used in the sinogram update.
  • D. Cross-Domain Iterative Training: The reprojected reference enters the sinogram update, yielding a next-iteration prior defined by the updated sinogram network.The consistency term is weighted by λiter.
  • D. Cross-Domain Iterative Training: Each outer iteration transfers information between projection and image domains, reducing prior error and promoting cross-domain consistency.The training and inference procedure repeats these updates for T outer iterations before final reconstruction.

E. Inference and Reconstruction

Inference combines the denoised sinogram prior with an observation residual before image-domain reconstruction, balancing noise suppression against detail preservation.

  • E. Inference and Reconstruction: The optimized sinogram network first estimates a prior from the noisy observation.This prior supplies the primary denoised estimate.
  • E. Inference and Reconstruction: Residual fusion reintroduces measurement information that the sinogram network may attenuate, with αinf controlling the suppression–detail trade-off.The fused sinogram is passed to the optimized image-domain network for final CT reconstruction.

IV. EXPERIMENTS

The evaluation covers experimental settings, comparisons on simulated and real low-dose CT data, and ablations of the framework’s main components.

  • IV. EXPERIMENTS: The experiments comprise settings, quantitative and qualitative comparisons, and ablation analyses.These components are reported in Sections IV-A through IV-C.

A. Experimental Settings

The study evaluates the method on simulated AAPM, LIDC-IDRI, and LoDoPaB-CT data at two dose levels, plus real cardiac and mouse CT data. Quantitative and qualitative comparisons assess denoising performance, artifact suppression, and structural preservation.

  • Real-data evaluation: Real-data evaluation included one GE Clinical Cardiac case and a mouse chest CT dataset acquired with specified clinical and micro-CT protocols.The GE case contained 984 projections over 360°; the mouse data used 2,000 projections over 360°.
  • Comparison methods: The comparisons included FBP, BM3D, supervised RED-CNN, and five self-supervised methods.The self-supervised baselines were N2V, N2N, B2U, N2Sim, and N2N-BS.
  • Evaluation criteria: PSNR, SSIM, and RMSE were reported as mean ± standard deviation, while qualitative assessment used matched windows, ROIs, and error maps.These visualizations were used to assess artifact suppression and structural preservation.
  • Quantitative results: The proposed method achieved the best PSNR, SSIM, and RMSE among self-supervised methods across all datasets and both dose levels.It outperformed supervised RED-CNN in 17 of 18 entries; AAPM SSIM at 1.0% dose was 95.72% versus 95.91%.
  • Qualitative results: Qualitatively, the method suppressed noise and streak artifacts while better preserving low-contrast boundaries, high-frequency structures, and tissue intensity transitions.It also produced more uniform error maps with fewer strong residual regions.
  • Real-data findings: Consistent performance on cardiac and mouse CT data indicates applicability to real low-dose CT data.The real-data figures include reconstructed images, local frequency spectra, and magnified ROIs.

C. Ablation Study

The ablation study examines image-domain refinement, posterior inference, variance matching, and cross-domain iteration. Results support the importance of these components and agreement between analytical and empirical variance-matching optima.

  • Component ablation: Removing the image-domain network caused the largest PSNR decrease and RMSE increase, highlighting the importance of image-domain refinement.The w/o Posterior variant removes latent count inference and explicit Gaussian noise separation, reducing the method to N2N-BS.
  • Component ablation: Removing posterior inference reduced the method to N2N-BS, while direct Poisson splitting was unsuitable for mixed LDCT noise in this setting.The ablation therefore tests the contribution of latent count inference and explicit Gaussian noise separation.
  • Cross-domain iteration: Removing cross-domain iteration degraded results after eliminating image-to-sinogram feedback and prior updates.The w/o Iteration variant retained sinogram-to-image processing but removed feedback in the reverse direction.
  • Variance matching: Reconstruction performance peaked near αvar_k = 0.7, while measured branch-to-observation variance ratios closely followed the theoretical curve.The empirical optimum and analytical matching point were both in the region of optimal reconstruction.
  • Variance matching: For balanced splitting, dk = 0.5 requires αvar_0.5 ≈ 0.707, supporting the variance-matching mechanism.The predicted matching point agreed with the observed optimum.

3) Effect of the Thinning Ratio:

The thinning-ratio study evaluates how branch fractions affect reconstruction and branch statistics, alongside iterative changes in prior error, residual correlation, and reconstruction quality.

  • Effect of the Thinning Ratio: Reconstruction performance was highest for the balanced split (0.5, 0.5) with equal branch fractions.Unequal splitting produced mismatched photon-count statistics and asymmetric branch predictions.
  • Effect of the Thinning Ratio: Unequal splitting yielded mismatched photon-count statistics and asymmetric branch predictions, supporting balanced branches.Together with the variance analysis, these results support the thinning and variance-matching design.
  • Cross-domain Iterative Refinement: During cross-domain iteration, normalized sinogram MAE and inter-branch residual correlation decreased in both sinogram and image domains.The reduction indicates weaker empirical dependence between regenerated branches as image-domain feedback refines the prior.
  • Cross-domain Iterative Refinement: PSNR and SSIM increased across iterations as prior refinement and pair quality improved.The reported trends link iterative prior refinement and regenerated-pair quality with improved reconstruction performance.

V. DISCUSSION

The work unifies physics-driven training-pair construction with cross-domain iteration, making pair generation dynamic and progressively refined through imaging physics and restoration feedback. Its main scope boundaries are model mismatch, approximation errors, and current two-dimensional processing.

  • Discussion: Physics-driven pair construction uses projection-domain noise formation while cross-domain restoration progressively refines the sinogram prior.This replaces static pair generation from fixed observations with a dynamically updated process.
  • Discussion: Under ideal Poisson–Gaussian modeling and exact posterior sampling, the generated branches satisfy strict conditional independence.Posterior approximation and prior-estimation errors can compromise this ideal property.
  • Discussion: When acquisition deviates substantially through effects such as scattering and beam hardening, residual correlations may remain between the branches.The paper identifies this model mismatch as the method’s main theoretical limitation.
  • Discussion: The current implementation is restricted to two-dimensional processing.Future work proposes three-dimensional or multi-slice extensions and broader acquisition models.

VI. CONCLUSION

The paper proposes self-supervised LDCT denoising that combines physics-driven pair construction with cross-domain iteration, and analyzes when the resulting branches are conditionally independent. The conclusion reports effectiveness for self-supervised LDCT denoising while the appendix formalizes the construction’s statistical properties.

  • Conclusion: The method constructs approximately independent training pairs from a single low-dose observation by integrating physics-driven pair construction with cross-domain iteration.The framework uses posterior sampling and stochastic splitting of Poisson and Gaussian components.
  • Conclusion: Experiments demonstrate the effectiveness and potential of the proposed method for self-supervised low-dose CT denoising.The conclusion states this outcome without restricting it to a particular dataset or metric.
  • Conclusion: Under the Poisson–Gaussian observation model and exact posterior sampling, the raw branches satisfy Y (1) ⊥Y (2) | y.Binomial thinning gives independent Poisson branches conditional on y, while Gaussian splitting and component independence support the joint result.
  • Conclusion: The branches have conditional means E[Y (k) | y] = dkµ and variances Var(Y (k) | y) = dk(µ + σ2), with d1 = α and d2 = 1 −α.These relations characterize the branch statistics before logarithmic transformation.
  • Conclusion: The first-order logarithmic approximation yields a variance-matching coefficient αvar k = √dk.The approximation relates transformed branch variance to the observation variance through the branch dose factors.
  • Conclusion: Theoretical analysis examines deviations in practical pair statistics using posterior approximation, prior-estimation errors, and total-variation analysis.The appendix introduces joint raw-branch distributions induced by practical and exact posteriors and bounds their discrepancy.
Loading 2609.02654v1…