Source-linked AI summary

Compressive Hyperspectral Imaging with Side Information

Xin Yuan, Tsung-Han Tsai, Ruoyu Zhu, Patrick Llull, David Brady, Lawrence Carin

arXiv:1502.06260v1cs.CV

TL;DR

The paper addresses hyperspectral reconstruction from highly compressed measurements, where a 3D datacube is mapped to a 2D image. It develops blind Bayesian dictionary learning with global-local shrinkage priors, extends it with RGB side information, and introduces an SLM-based camera. Experiments on synthetic and real measurements report high-quality reconstruction, camera feasibility, and benefits from side information.

  • Problem

    Hyperspectral reconstruction from a single compressed measurement is challenging because the 3D datacube is mapped to a 2D observation, while the dictionary needed for blind reconstruction is unknown.

  • Method

    The method learns a dictionary in situ with global-local shrinkage priors and can learn a coupled dictionary from compressed measurements and corresponding RGB images; an SLM camera performs spectral coding.

  • Results

    Experiments on synthetic data and real CASSI and SLM-CASSI measurements demonstrate high-quality inversion, camera feasibility, and improved reconstruction with RGB side information.

  • Takeaways & Limitations

    RGB side information can substantially improve blind compressive hyperspectral reconstruction, and the framework can extend to other dictionary-learning approaches.

Abstract

from arXiv · show

A blind compressive sensing algorithm is proposed to reconstruct hyperspectral images from spectrally-compressed measurements.The wavelength-dependent data are coded and then superposed, mapping the three-dimensional hyperspectral datacube to a two-dimensional image. The inversion algorithm learns a dictionary {\em in situ} from the measurements via global-local shrinkage priors. By using RGB images as side information of the compressive sensing system, the proposed approach is extended to learn a coupled dictionary from the joint dataset of the compressed measurements and the corresponding RGB images, to improve reconstruction quality. A prototype camera is built using a liquid-crystal-on-silicon modulator. Experimental reconstructions of hyperspectral datacubes from both simulated and real compressed measurements demonstrate the efficacy of the proposed inversion algorithm, the feasibility of the camera and the benefit of side information.

I. INTRODUCTION

CASSI compressively encodes a hyperspectral datacube into a 2D measurement, making reconstruction underdetermined and motivating blind dictionary learning. The paper adds global-local shrinkage priors, RGB side information, and an SLM-based camera to improve reconstruction and provide an alternative coding design.

  • CASSI imaging system: CASSI codes each spectral channel with a distinct 2D pattern, spectrally shifts the coded images, and multiplexes them onto a monochrome detector.The disperser provides wavelength-dependent shifts, giving channels unique coding patterns.
  • CASSI imaging system: A single 2D CASSI measurement can support a 3D spectral estimate, but the resulting forward model is highly underdetermined.Recovery therefore requires compressive-sensing inversion algorithms.
  • Blind reconstruction: Blind compressive sensing jointly learns a dictionary from measurements and reconstructs hyperspectral patches as sparse combinations of shared dictionary atoms.The proposed model replaces explicit sparsity enforcement with compressibility induced by global-local shrinkage priors.
  • Side information: RGB side information enables a coupled dictionary learned from one CASSI measurement and the corresponding RGB image, improving reconstruction quality.The RGB image is strongly correlated with the hyperspectral image and can be acquired with an off-the-shelf color camera.
  • Camera design: The proposed SLM camera jointly modulates spatial and spectral information on the intermediate image plane without requiring a dispersive element.This design differs from CASSI and DMD-modulated compressive imagers in its multiplexing mechanism.

B. A New Camera: SLM-CASSI

SLM-CASSI uses wavelength-dependent spatial modulation to multiplex hyperspectral information onto a 2D detector without a dispersive element. The prototype requires calibrated transmission patterns because projection errors affect reconstruction.

  • SLM-CASSI operation: SLM-CASSI encodes 3D spatial-spectral information onto a 2D grayscale detector using wavelength-dependent transmission patterns.The SLM’s birefringent modulation changes polarization and produces spectral transmission codes for compressive measurement.
  • SLM-CASSI operation: Unlike conventional CASSI and DMD-based imagers, the proposed camera multiplexes spectral channels without a dispersive element.The SLM jointly modulates spatial and spectral information on the intermediate image plane.
  • Experimental setup: The prototype combines relay and imaging optics, polarization components, an LCoS SLM, and a monochrome CCD detector.The detector has 2048×2048 pixels of 7.4 µm pitch, while the SLM has a 1920×1080 active area with 8 µm pixel pitch.
  • System calibration: Calibration measures spectral transmission responses to construct a system operator containing the SLM’s possible coding patterns.Monochromatic illumination is quantized into bands with 7.5–8 nm FWHM for improved response representation.
  • System calibration: Within 450–680 nm, 8-bit SLM voltage variation produces transmission ratios from 0.08 to 0.96.The voltage- and wavelength-dependent response supplies the grayscale modulation used by the camera.
  • System calibration: Optical aberrations and detector–SLM alignment discrepancies can make the ideal one-to-one projection inaccurate, requiring careful calibration.The authors report that adaptive sensing contributed little because of mismatch between designed codes and calibration, while random coding performed well.
  • Shared mathematical model: Both CASSI and SLM-CASSI reconstruct Nλ spectral images from one measurement, giving a compression ratio of Nλ.Their measurements are represented as sums of coded spectral slices, with SLM-CASSI patterns determined by the SLM response.

III. RECONSTRUCT HYPERSPECTRAL IMAGES WITH BLIND COMPRESSIVE SENSING

The paper develops a blind Bayesian compressive-sensing model that jointly learns a dictionary and sparse patch coefficients from CASSI or SLM-CASSI measurements. Global-local shrinkage promotes compressibility, while patch-specific forward operators incorporate random masks and calibration.

  • Model overview: The proposed blind Bayesian CS model reconstructs hyperspectral images from CASSI or SLM-CASSI measurements and also generalizes to denoising and inpainting.The model learns the representation from the measurements rather than requiring a dictionary a priori.
  • Dictionary learning model: The datacube is divided into vectorized 3D patches, each represented as a dictionary times a coefficient vector.For patch n, x_n is modeled using dictionary D and coefficients s_n, while the corresponding measurement patch is included in the observation model.
  • Measurement model: The measurement model includes additive Gaussian noise with precision α_0 and patch-specific forward operators Ψ_n.Each Ψ_n reflects the random mask and system calibration associated with that patch.
  • Model extensions: The model can represent spatially varying noise by imposing different noise models on different patches.This extension is stated as straightforward for non-uniform denoising problems.
  • Shrinkage prior: The global-local shrinkage prior encourages most dictionary coefficients to be small, imposing compressibility on hyperspectral patches.The global parameter τ_n scales all coefficients in a patch, while local weights Φ_k,n control individual coefficients.
  • Dictionary size control: A multiplicative gamma prior increasingly shrinks later dictionary atoms and supports automatic removal of atoms with small weights.The resulting ordering of atom weights helps infer dictionary importance and reduce overfitting.

B. The Statistical Model

The statistical model combines data fitting, dictionary regularization, adaptive weighting, and global-local shrinkage within a Bayesian posterior over the unknown parameters.

  • Prior specification: Broad hyperpriors are assigned by setting a_0 through h_0 to 10^-6, and coefficients are scaled by the noise precision α_0.These settings define the stated prior specification for the full statistical model.
  • Posterior formulation: The model infers Θ = {D, S, Λ, ϵ} through the log joint posterior.These parameters comprise the dictionary, coefficients, shrinkage-related variables, and noise terms.
  • Posterior formulation: The posterior includes an ℓ2 data-fit term, ℓ2 dictionary regularization, and a weighting parameter α_0 that balances them.The weighting is updated during inference rather than fixed in advance.
  • Shrinkage structure: Global and local shrinkage parameters regularize the coefficients to promote compressibility in high-dimensional hyperspectral patches.The model uses τ_n and local shrinkage parameters rather than imposing only direct sparsity.

C. Related Models

The paper replaces conventional sparsity modeling with global-local shrinkage priors in a blind dictionary-learning framework. Bayesian inference estimates model parameters without cross-validation, while the framework also supports related restoration tasks.

  • Shrinkage modeling: Global-local shrinkage priors impose compressibility on sparse codes through global and local shrinkage parameters.The model also applies ℓ2 regularization to the sparse codes and dictionary atoms.
  • Bayesian inference: Unlike ℓ1-based alternatives, the Bayesian model infers shrinkage and regularization parameters through posterior distributions rather than setting them by hand.The approach avoids cross-validation for these parameter choices.
  • Model relation: Replacing the shrinkage prior with a spike-slab prior reduces the model to BPFA, whereas the proposed prior is intended for high-dimensional hyperspectral patches.The paper distinguishes direct sparsity from the proposed compressibility prior.
  • Inference: Local conjugacy yields closed-form conditional posteriors, enabling straightforward Gibbs-sampling MCMC inference.A variational Bayes alternative is also available by replacing conditioning variables with their moments.
  • Broader applicability: The framework is presented as applicable beyond hyperspectral reconstruction, including denoising and inpainting.Benchmark color-image experiments compared the method with K-SVD and BPFA, and the proposed algorithm constantly performed better.

E. RGB Images as Side Information

RGB images from the same scene are incorporated as side information alongside a single CASSI measurement. Joint dictionary learning uses the paired data to improve hyperspectral reconstruction while retaining a general formulation across cameras.

  • Motivation: An RGB image can be captured by an off-the-shelf camera in the unused optical path and used to aid reconstruction from a single compressed measurement.The RGB and hyperspectral images depict the same scene.
  • Measurement model: With an additional RGB camera, the joint dataset has compression ratio (Nλ + 3) : 4.The measurement combines the RGB image with the CASSI measurement.
  • Joint dictionary learning: The algorithm learns a super-dictionary and sparse codes jointly from paired RGB and compressed hyperspectral patches.The top hyperspectral rows are selected, overlapping patches are averaged, and the result is reformatted into a 3D image.
  • Reconstruction benefit: Because RGB images are fully observed and share the scene with the hyperspectral data, they facilitate dictionary learning and improve reconstruction quality.The learned dictionary embeds both wavelength-specific images and pixel spectra.
  • Design choice: The formulation is designed as a general and robust way to combine RGB images with CASSI measurements despite camera-dependent sensor quantum efficiencies.The experiments use RGB side information measured separately by an additional camera.

IV. EXPERIMENTAL RESULTS

Experiments evaluate the blind CS method on simulated and real compressed hyperspectral measurements using reconstruction quality and spectral-correlation metrics. The proposed shrinkage method performs best when RGB side information is available, while runtime remains comparable to selected alternatives.

  • Experimental setup: The experiments use single compressed measurements and compare the proposed blind CS method with TwIST and linearized Bregman.The evaluation uses PSNR across wavelengths and correlation between reconstructed and reference spectra.
  • Experimental setup: The implementation uses 8×8 spatial patches, typically sets K = 64 dictionary atoms, and runs on a 3.3GHz desktop with 16GB RAM.Increasing K beyond 64 did not significantly change the outcome in the reported experiments.
  • Runtime: The proposed model has computational time similar to BPFA and linearized Bregman, while TwIST is faster but generally produces worse reported metrics.Linearized Bregman and TwIST require about 20 and 14 minutes, respectively, for the 768×1024×24 bird data.
  • Scope: The experiments focus on single-measurement reconstruction, whereas the proposed model can also be used in multiple-measurement scenarios.The cited prior method did not show operation with a single measurement, and the paper reports worse results when it was tried.
  • Nature scenes: The proposed shrinkage method outperforms the other algorithms on natural-scene hyperspectral data, especially when RGB images are used as side information.The dataset contains 33 channels spanning 400–720nm at 10nm intervals.

2) Bird data:

Bird-data experiments show that RGB side information substantially improves the proposed shrinkage method’s hyperspectral reconstruction, especially in spectral accuracy and image clarity. Across real datasets and camera configurations, the method generally outperforms or matches competing approaches.

  • Bird data: The shrinkage method coupled with RGB side information provides the best reconstruction PSNR and spectral matches on the 24-channel bird data.Without RGB, reconstructions appear noisy and spectra are poorly represented; with RGB, spectral matches are comparable to TwIST while preserving more detail.
  • Bird data: On original-CASSI real data, the proposed method without side information provides spectra similar to TwIST, while linearized Bregman performs poorly.The RGB image was not used because it was not well aligned with the CASSI measurement.
  • Bird data: RGB side information produces clearer images and more accurate spectra than the proposed model without RGB on aligned multiframe-CASSI bird data.The RGB-assisted reconstruction is reported as best for both image clarity and spectral accuracy.
  • Bird data: The proposed method outperforms TwIST on the M&M dataset captured by the SLM-CASSI camera without RGB side information.The dataset contains 30 spectral channels and uses fiber-optic spectrometer measurements as reference spectra.
  • Bird data: The berry reconstruction separates green leaf responses at 520∼590nm from red berry responses at 610∼680nm.This demonstrates spectral localization of the scene’s leaf and berry components.

V. CONCLUSIONS

The paper concludes that Bayesian blind compressive sensing with global-local shrinkage priors enables high-quality reconstruction from real compressed hyperspectral measurements. RGB side information improves reconstruction, while the SLM-based camera demonstrates a complementary hardware design with distinct trade-offs.

  • V. CONCLUSIONS: The Bayesian dictionary-learning model enables high-quality inversion of compressed hyperspectral images captured by real cameras.Global-local shrinkage priors impose compressibility on dictionary coefficients to extract more information from dictionary atoms.
  • V. CONCLUSIONS: Integrating compressed measurements with RGB images significantly improves hyperspectral reconstruction quality under the blind CS framework.The conclusion identifies side information as a central source of the reconstruction improvement.
  • V. CONCLUSIONS: The SLM-CASSI camera demonstrates the feasibility of spectral coding with an SLM and supports flexible masking and a separate RGB camera.The SLM refresh rate also offers potential for video-rate compressive sensing and multiframes adaptive sensing.
  • V. CONCLUSIONS: The original coded-aperture/disperser architecture is more reliable, smaller, and less expensive than SLM-CASSI, but lacks coding flexibility and separate-camera side information.The conclusion frames the camera designs as a trade-off between compact reliability and coding or side-information flexibility.
  • V. CONCLUSIONS: The side-information framework may also benefit Gaussian-mixture-model dictionary learning approaches that learn unions of subspaces.This is presented as a possible generalization beyond the proposed blind CS model.

A. MCMC Inference

The MCMC inference section specifies sampling steps for latent variables in the Bayesian model. It introduces inverse-Gaussian and generalized inverse-Gaussian distributions used in those updates.

  • A. MCMC Inference: The MCMC procedure includes a sampling step for αk,n associated with an inverse-Gaussian distribution.The passage identifies IG as the inverse-Gaussian distribution.
  • A. MCMC Inference: The procedure samples Φk,n using a generalized inverse-Gaussian distribution.The generalized inverse-Gaussian distribution is denoted GIG(x : a, b, p).
  • A. MCMC Inference: The generalized inverse-Gaussian formulation uses Kp(θ), the modified Bessel function of the second kind.This function appears as part of the distributional formulation used in the MCMC updates.

B. Variational Bayesian Inference

The variational Bayesian inference section replaces full posterior inference with a factorized approximation and derives coordinate-wise updates. The approximation is optimized through the evidence lower bound, with parameters updated from the model posterior.

  • B. Variational Bayesian Inference: Mean-field variational Bayes approximates the posterior p(Θ|Y) with a simpler distribution q(Θ) to improve runtime.The observed data are represented by Y, while Θ denotes the independent latent variables.
  • B. Variational Bayesian Inference: The model defines Θ as the collection {D, S, Λ, ϵ} and assumes complete factorization across latent variables.The factorization is written as q(Θ) = ∏i qi(Θi).
  • B. Variational Bayesian Inference: Optimizing q by minimizing KL divergence between the true and approximate posteriors estimates the conditional posterior distribution.The variational objective is expressed through the Kullback-Leibler divergence.
  • B. Variational Bayesian Inference: Because ln p(Y) is fixed with respect to q(Θ), maximizing the ELBO is equivalent to minimizing KL divergence.This establishes the objective used to obtain the variational approximation.
  • B. Variational Bayesian Inference: Under the factorized model, each variational parameter is independently updated using coordinate-wise inference equations.The update equations follow from the posterior distributions used in Gibbs sampling, with expectations replacing sampled variables.

VII. MORE RESULTS

Benchmark color-image experiments evaluate the proposed model on inpainting and denoising across missing-data ratios and Gaussian-noise levels. The proposed algorithm achieves the highest reported PSNR and requires relatively few inference iterations.

  • Benchmark Images for Inpainting and Denoising: The proposed algorithm consistently provides the highest PSNR for inpainting and denoising benchmark color images.Inpainting varies observed pixel percentages, while denoising varies the standard deviation σ of additive zero-mean Gaussian noise.
  • Benchmark Images for Inpainting and Denoising: 20 iterations are used for denoising and 100 iterations for inpainting, at approximately 1 second per iteration on a 256 × 256 × 3 image.These timings use an i5 CPU with nonoptimized MATLAB code.
  • Benchmark Images for Inpainting and Denoising: The noisy-image dictionary example uses K = 256 atoms, of which 147 are significant.The example uses σ = 5 and patch size 7 × 7 × 3; inferred |νk| values are sorted by absolute magnitude.
  • Benchmark Images for Inpainting and Denoising: The corrupted and restored-image figures organize each row by image and each column by observed-data ratio or noise level.For denoising, the noise level is defined by the standard deviation σ of additive zero-mean Gaussian noise.
Loading 1502.06260v1…