Source-linked AI summary

Compact Snapshot Spectral Imaging with Calibration-Free Aperture Diffraction

Tao Lv, Quan Yuan, Shiqiao Li, Chenglong Huang, Linsen Chen, Chongde Zi, Shuming Wang, Xun Cao

arXiv:2608.29230v1cs.CV

TL;DR

Existing SSI systems are constrained by large form factors, repeated calibration, and persistent simulation-to-reality gaps. The paper proposes ADIS, combining orthogonal diffraction optics and theoretically computed PSFs with ODAUVST reconstruction. ADIS achieves calibration-free full-resolution SSI in a standard RGB-camera footprint, while the authors identify reduced light throughput as a limitation.

  • Problem

    Traditional SSI systems are large and require repeated calibration, while miniaturized DOE systems retain calibration needs and simulation-to-reality gaps that can introduce artifacts.

  • Method

    ADIS combines a binary orthogonal diffraction lens with a sensor and theoretically computed PSFs, while ODAUVST integrates ODAUF with VST for spectral reconstruction.

  • Results

    ADIS achieves full-resolution calibration-free SSI within a standard RGB-camera footprint; ODAUVST reaches 33.31dB PSNR and 0.951 SSIM on ADIS spectral reconstruction.

  • Takeaways & Limitations

    The prototype and experiments support direct migration from theoretically computed PSFs to real-world acquisition while tolerating lens-dependent optical variations.

  • Takeaways & Limitations

    Amplitude diffraction coding reduces overall light throughput, and improving throughput without degrading performance remains an open issue.

Abstract

from arXiv · show

Snapshot Spectral Imaging (SSI) provides high-dimensional temporal-spatial-spectral observation to uncover intrinsic physical characteristics. However, its complex system and repetitive calibration requirements hinder edge applications. Here, we propose a compact, cost-effective, calibration-free SSI method, Aperture Diffraction Imaging Spectrometer (ADIS), which consists only of a diffractive lens with a binary mask and a Bayer-filtered sensor, requiring no additional physical footprint compared to standard RGB cameras. ADIS disperses and multiplexes wavelengths, mapping energy to distinct sensor locations, enabling full-resolution recovery from superpixel-level encodings. ADIS directly leverages theoretically computed PSFs to enable calibration-free spectral reconstruction, while tolerating lens-dependent variations across different optical configurations and bridging the gap between simulation and reality. To achieve SSI by solving a sparsely-constrained inverse problem, we introduce the Orthogonal Diffraction-Aware Unfolding Framework (ODAUF) with Voxel Shift Transformer (VST) for improved orthogonal diffraction perception. Integrating VST into ODAUF forms the efficient Orthogonal Diffraction-Aware Unfolding Voxel Shift Transformer (ODAUVST), delivering excellent recovery and reduced parameters. By elaborating on theory, systematic and comprehensive comparing, and demonstrating real SSI results, we validate the superiority of ADIS, achieving calibration-free full-resolution SSI within a commercial camera footprint.

I. INTRODUCTION

SSI offers rich spectral information but traditional systems are large, complex, and calibration-intensive, limiting edge deployment. ADIS addresses these constraints with compact orthogonal diffraction encoding, theoretically computed PSFs, and reconstruction methods designed for calibration-free full-resolution SSI.

  • Motivation: Hyperspectral imaging distinguishes materials with similar colors by capturing fine-grained spectral data at each spatial location.
  • Limitations of existing SSI: Traditional SSI systems require dispersive optics, masks, relay lenses, and imaging lenses, producing large form factors and demanding repeated calibration.
  • Limitations of existing SSI: DOE-based miniaturized systems still require calibration, while simulation-to-reality gaps can introduce artifacts in reconstructed outputs.
  • ADIS: ADIS combines a binary orthogonal mask with a lens and sensor to provide compact SSI without increasing the camera’s physical footprint or requiring calibration.The ODL is physically characterized and accommodates lens-dependent convolutional degradation across optical configurations.
  • ADIS: Orthogonal diffraction produces lattice-like PSFs that multiplex wavelengths onto distinct sensor locations, enabling full-resolution recovery from superpixel-level filtered measurements.Complementary masks can generate similar diffraction patterns while improving energy efficiency according to Babinet’s principle.
  • Validation: The study validates ADIS through optical analysis, algorithm comparisons, hardware design comparisons, and single-exposure real spectral imaging results.The reported outcome is full-resolution SSI without system calibration within the footprint of standard RGB cameras.
  • Reconstruction: ODAUF uses diffraction-projection guidance and spatial-spectral priors to mitigate the ill-posed reconstruction problem, while VST improves degradation perception and generalizability.Integrating VST into ODAUF yields ODAUVST for high-resolution simulations and real experiments, with excellent restoration and reduced parameters.

B. Array-Filtered Encoding Methods

Array-filtered and aperture-encoding approaches compact spectral sensing through tiled filters or aperture-plane dispersive optics, but face resolution, complexity, and calibration constraints. ADIS instead combines compact optical encoding with computational reconstruction and analyzes diffraction-based imaging behavior.

  • Array-Filtered Encoding Methods: Array-filtered encoding uses tiled spectral filter arrays, but increasing sampled channels reduces spatial resolution.
  • Aperture Encoding Methods: Aperture encoding uses dispersive optics at the aperture plane to spectrally encode the scene.
  • Aperture Encoding Methods: CTIS combines differently oriented gratings to measure linear projections, but its occlusion mask and spectral projections require a complex structure and repetitive calibration.
  • ADIS Configuration: ADIS uses a diffractive lens with an orthogonally distributed binary mask and a commercial RGB camera, while computational decoding reconstructs hyperspectral images from compressed measurements.
  • Diffraction Model: The derivation concludes that ADIS is depth-invariant within a specific depth range under the approximation ξ ≈ Z.
  • Diffraction Model: The diffraction model separates the pattern into a rectangular-aperture diffraction factor D(x,y,λ) and a multi-slit interference factor P(x,y,λ).

C. Hardware Parameter Selection and Logic

Hardware selection balances spectral resolution, diffraction strength, inverse-problem difficulty, and reconstruction fidelity. The analysis selects mask periods and focal-length-to-period ratios that provide sufficient encoding with stable reconstruction.

  • Constraint from Target Spectral Resolution: For 28 spectral bands across 450−650nm, the third-order dispersed stripe must span at least 28 pixels, requiring d < 310.56µm.
  • Trade-off between Inverse Problem Solving and Reconstruction Fidelity: Smaller mask periods strengthen diffraction and separate spectral information more explicitly, but make inverse reconstruction harder; larger periods simplify inversion while reducing spectral fidelity.
  • Trade-off between Inverse Problem Solving and Reconstruction Fidelity: For f = 50mm and spixel = 3.45µm, mask periods d = 180−310µm consistently yield high spectral reconstruction fidelity.
  • Trade-off between Inverse Problem Solving and Reconstruction Fidelity: The prototype adopts d = 200µm as a calculation-friendly and fabrication-friendly mask period.
  • Generalized Parameter Selection: For 450−650nm systems, the recommended effective diffraction-strength ratio is f/d = 161.29−277.78.

IV. SNAPSHOT SPECTRAL IMAGING WITH APERTURE DIFFRACTION

ADIS models hyperspectral capture through Bayer-filter responses and wavelength-dependent diffraction PSFs, then reconstructs spectra from the resulting multiplexed measurements. Its real-lens model accounts for lens-dependent convolutional degradation, while reconstruction outputs reflect the employed lens’s finite imaging capability.

  • Forward Formation Model: ADIS forms Bayer measurements by integrating spectrally filtered convolutions of hyperspectral channels with wavelength-dependent PSFs.The forward model uses sensor spectral sensitivity and convolution with wavelength-specific PSFs.
  • Forward Formation Model: The discrete model represents the measurement as y = Φx + n, where Φ combines sensor sensitivity and PSF convolution determined by binary masks.The hyperspectral image vector contains Λ wavelength channels, while Φ maps it to the Bayer measurement space.
  • Calibration-Free Reconstruction: Real-lens PSFs are modeled as simulated PSFs convolved with wavelength-dependent lens degradation caused by diffraction, aberrations, and other optical imperfections.This formulation captures the difference between ideal thin-lens simulations and practical lens behavior.
  • Calibration-Free Reconstruction: Networks trained with simulated PSFs reconstruct the hyperspectral signal convolved with lens degradation, so the output represents SSI within the employed lens’s finite imaging capability.The reconstructed target is therefore interpreted under the real lens’s imaging limitations rather than as an ideal unconstrained signal.

C. Orthogonal Diffraction-Aware Unfolding Framework

ODAUF unfolds an optimization procedure for ADIS reconstruction, combining data fidelity, regularization, parameter estimation, linear updates, and denoising. Its linear projection efficiently incorporates the sensing model, while learned denoisers implement the regularization step.

  • Optimization Analysis: ODAUF formulates ADIS reconstruction as an optimization problem combining forward-model data fidelity with a regularization prior.The prior may represent sparsity, total variation, low-rank structure, or a network-based image prior.
  • Optimization Analysis: HQS introduces an auxiliary variable z, converting the constrained reconstruction into alternating subproblems for the image estimate x and regularized variable z.The penalty term encourages x and z to approach one another during iterative optimization.
  • ODAUF: The x-update becomes an efficient linear projection using a gradient-descent step instead of directly evaluating the computationally prohibitive closed-form inverse.The framework can iteratively increase the penalty parameter ρ, balancing agreement between x and z against convergence speed.
  • ODAUF: The z-update is interpreted as Gaussian denoising, with iteration-specific parameters controlling the denoising level in each unfolding stage.ODAUF therefore treats denoisers as learnable regularization terms within the unfolded optimization process.
  • ODAUF: ODAUF uses a hyperparameter estimator Θ, a linear projection Ξ, and a denoiser D to form its unfolded reconstruction stages.Θ receives the compressed measurement and sensing matrix, Ξ updates x using the ADIS model, and D solves the denoising subproblem.

D. Realization of ODAUF

ODAUF embeds ADIS imaging and hardware priors into an unfolding reconstruction pipeline, while VST adds adaptive voxel-shift denoising and attention for diffraction-aware recovery.

  • ODAUF: ODAUF addresses consistent diffractive aliasing by modeling ADIS degradation and incorporating downsampled PSF and filter representations into reconstruction.The framework downsamples PSF spatial resolution and filter channels to 256 × 256 × 3 for network processing.
  • VST: VST is an adaptive degradation-aware denoiser embedded in ODAUF that extracts diffraction features and high-dimensional information through voxel shifts.Its three-layer U-shaped design uses integer and fractional shifts, channel shuffling, and attention for global and long- and short-range modeling.
  • Initialization: VST initialization incorporates downsampled hardware priors, filter functions, and measurements before reconstruction.The priors are reduced to match the network representation and concatenated with the measurement before upsampling.
  • VSAB and VS-MSA: VSAB combines layer normalization, VS-MSA, and an FFN, with VS-MSA projecting features into queries, keys, and values before shift-aware self-attention.The first stage learns fractional shifts from Q1, K1, and V1 through LSM; later stages split channels and use windowed and shuffled attention.
  • LSM and CSM: LSM extracts fractional shift parameters, while CSM combines fixed integer shifts with learnable continuous shifts implemented through differentiable grid-based interpolation.The learned fractional shift is computed as ς = ω · tanh(ζ), with bilinear interpolation and reflection padding used during warping.

V. SIMULATION EXPERIMENT ANALYSIS

The simulation experiments construct ADIS measurements from rendered hyperspectral data and evaluate reconstruction on CAVE-1024 and KAIST scenes.

  • Simulation Setup: The experiments extract a 256 × 256 × 28 spectral region, apply a matching filter, and integrate along the spectral dimension to generate two-dimensional measurements.This image-formation model is used to maintain consistency when preparing rendered measurements for reconstruction.

A. Comparison with Other Reconstructions

ADIS reconstruction is compared with compact DOE-based and prism-based spectral imaging systems using simulated ground-truth spectral images.

  • System Comparison: ADIS is evaluated against a DOE-based system and a prism-based system using ten ground-truth spectral images.
  • System Comparison: The comparison assesses compact computational spectral imaging methods across reconstructed spectral images from the three systems.

1) Quantitative Comparisons:

Across algorithmic, system-level, framework, qualitative, and noise evaluations, ODAUVST and ADIS achieve strong reconstruction quality with compact, calibration-free operation and improved robustness.

  • Quantitative Comparisons: 33.31 dB PSNR and 0.951 SSIM are achieved by ODAUVST, which exceeds CSST-9stg by 0.76 dB and Restormer by 1.36 dB.It uses 29.96% of Restormer’s parameters and 61.00% of its FLOPS, and 69.05% of CSST-9stg’s parameters and 76.09% of its FLOPS.
  • Qualitative Comparisons: ODAUVST-5stg produces sharper textures, finer details, and spectral profiles more closely matching ground truth than the compared reconstruction methods.Alternative ADIS reconstructions show over-smoothing, chromatic artifacts, and speckle patterns absent from the ground truth.
  • System Comparisons: ADIS achieves superior spatial and spectral reconstruction accuracy relative to DOE-based and prism-based systems while using a physically describable binary mask without calibration.The binary mask adds no footprint or operational complexity according to the comparison passage.
  • Unfolding Framework Comparisons: ODAUF improves reconstruction without significantly increasing memory or computational costs, and paired with VST it surpasses prior methods in imaging quality.
  • Noise Robustness: ADIS remains more stable than the DOE-based system under noise, whose reconstruction quality degrades markedly even under mild noise.The reported comparison uses additive Gaussian noise with standard deviation ϖ = 0.01.

VI. REAL EXPERIMENTAL ANALYSIS

ADIS prototypes demonstrate calibration-free, full-resolution real-world spectral imaging using computed PSFs, including ColorChecker accuracy, spatial detail preservation, metamer discrimination, and dynamic capture.

  • Prototype systems: The ADIS-V1 and ADIS-V2 prototypes use orthogonal masks attached to 50mm lenses, with ADIS-V2 employing a commercial compound lens.Both systems reuse PSFs simulated for the 50mm singlet, while a 450–650nm filter and adjustable iris constrain operation.
  • Real reconstruction: 2448 × 2048 × 28 spectral reconstructions preserve sharp textures and well-defined spatial content with minimal artifacts in real scenes.The corresponding measurements are captured at 2448×2048 spatial resolution.
  • Spectral accuracy: ColorChecker spectra closely match point-spectrometer measurements, achieving spectral resolution below 8nm with the current configuration.The reconstruction retains well-defined structures and minimal artifacts.
  • Spatial resolution: ADIS restores high-frequency spatial details, with post-reconstruction modulation transfer functions demonstrating improved spatial accuracy across textured spectral bands.The evaluation compares ISO12233 and highly textured scenes before and after spectral reconstruction.
  • Application demonstrations: ADIS distinguishes metameric materials with similar RGB appearances by recovering spectra that closely agree with scanning-spectral-camera references.The system thereby preserves discriminative spectral information beyond RGB imaging.
  • Application demonstrations: ADIS-V2 captures autonomous-driving spectral video at 25 FPS with good spatio-temporal consistency throughout the sequence.The experiments cover dynamic outdoor scenes in addition to static natural scenes and material discrimination.

VII. ABLATION STUDY AND HARDWARE ANALYSIS

Ablations show that learnable fractional shifts, reflection padding, ODAUF, and VST improve or preserve reconstruction efficiency, while square masks offer the most practical hardware choice.

  • LSM ablation: A learnable fractional-shift range improves full-resolution generalization over a fixed range that can overfit a limited test range.Fixed integer steps of perform best when ω is either fixed or learnable.
  • CSM ablation: Reflection padding delivers the best reconstruction performance by avoiding boundary artifacts and preserving edge continuity.The comparison includes reflection, zero, and border padding.
  • ODAUF and VST ablation: 1.73dB in PSNR and 0.015 in SSIM are gained by integrating ODAUF and VST over the COPF-and-SST combination at K = 5.Standalone ODAUF and VST also improve PSNR and SSIM relative to CSST.
  • ODAUF and VST ablation: 0.76dB in PSNR and 0.015 in SSIM favor ODAUVST-5stg over CSST-9stg while using 69.05% of parameters and 76.09% FLOPS.The result supports more efficient reconstruction with fewer stages.
  • Mask-form analysis: Triangular and pentagram masks achieve marginally higher simulated accuracy, but square-hole arrays remain more practical because asymmetric masks complicate alignment and fabrication.Circular masks are slightly less effective than square-hole masks in empirical tests.
  • Hardware analysis: PSF depth variance is negligible beyond 3.0m when the system is focused at infinity, while close-range SSI requires refocusing at the target depth.The result indicates that far-field PSFs primarily depend on wavelength.

B. Lower Simulation-to-Reality Gap

ADIS reduces simulation-to-reality mismatch through simple binary-mask encoding, theoretically computed PSFs, and tolerance to optical and fabrication variations, while retaining broad scene performance.

  • Simulation-to-reality gap: ADIS uses a simple binary mask and theoretically computed PSFs to support calibration-free reconstruction despite lens-dependent optical variations.The approach is described as providing a more seamless transition from simulation to real acquisition.
  • Scene robustness: ODAU-VST reconstructions remain high quality in both edge-rich and smooth-content scenes, with learnable fractional shifts improving real-world generalization.The method reduces edge dependence from algorithmic and coding perspectives.
  • Optical variation: PSF lattice-point positions remain nearly identical across compound, cemented achromatic, and biconvex lenses sharing the same 50mm effective focal length.The comparison links similar diffraction propagation distances to consistent PSF structure.
  • Fabrication tolerance: Random mask errors of 2µm and 5µm preserve high PSF structural consistency for a 100µm line width and 200µm period design.The masks are therefore characterized as tolerant of realistic manufacturing deviations.
  • Fabrication tolerance: ADIS masks combine fabrication errors below 0.2µm with tolerance to residual deviations, supporting calibration-free hardware operation.Standard laser lithography can reproduce the designed mask with high fidelity.
  • Limitations: Amplitude coding reduces overall light throughput, while complementary masks can retain 75% throughput but may increase filtering reliance and compromise spectral fidelity in unknown scenes.Improving throughput without degrading performance remains an open issue.
  • Limitations: Complex or high-frequency illumination spectra can distort reconstructed spectra because CAVE and KAIST training data use reflectance data, with limited mitigation under mixed lighting.Dot-multiplying the light-source spectrum with the dataset can help in specific scenarios.

X. BIOGRAPHY SECTION

The biography section lists the authors’ academic affiliations, degrees, positions, and research interests, primarily at Nanjing University.

  • Author biographies: Tao Lv, Quan Yuan, Shiqiao Li, Chenglong Huang, Linsen Chen, Chongde Zi, Shuming Wang, and Xun Cao are affiliated with Nanjing University.The biographies describe their academic backgrounds and research areas.
  • Author biographies: The authors’ research interests span computational photography, spectral imaging, optics, nanophotonics, metasurfaces, and computer vision.Several biographies specifically mention computational imaging and spectral reconstruction.
  • Author biographies: Shuming Wang specializes in nanophotonics, metasurfaces, plasmonics, and quantum optics, while Xun Cao is identified as a professor and IEEE member.Their biographies include prior academic and visiting research appointments.
Loading 2608.29230v1…