Source-linked AI summary
DiffuserCam: Lensless Single-exposure 3D Imaging
Nick Antipa, Grace Kuo, Reinhard Heckel, Ben Mildenhall, Emrah Bostan, Ren Ng, Laura Waller
TL;DR
DiffuserCam addresses the challenge of compact, fast, high-resolution single-shot 3D imaging by encoding a volume with a diffuser and recovering it computationally. Its compressed-sensing reconstruction reaches 100 million voxels from a single 1.3 megapixel image, while the paper shows that effective resolution depends on scene content.
Problem
Existing 3D imagers are often bulky, scanning or multi-shot systems trade speed for resolution, and existing single-shot methods have low resolution.
Method
DiffuserCam uses a thin phase diffuser to create position-dependent pseudorandom caustic patterns and reconstructs sparse 3D intensity from one 2D measurement using a physical model and calibration.
Results
100 million non-uniformly spaced voxels are reconstructed from a single 1.3 megapixel image using a commodity-hardware prototype.
Takeaways & Limitations
The system demonstrates compact single-shot lensless volumetric imaging with true depth sectioning, while its nonlinear reconstruction makes performance object-dependent.
Takeaways & Limitations
Effective resolution varies with scene content, and standard two-point resolution can represent only a best-case scenario for this nonlinear system.
Abstract
from arXiv · showhide
We demonstrate a compact and easy-to-build computational camera for single-shot 3D imaging. Our lensless system consists solely of a diffuser placed in front of a standard image sensor. Every point within the volumetric field-of-view projects a unique pseudorandom pattern of caustics on the sensor. By using a physical approximation and simple calibration scheme, we solve the large-scale inverse problem in a computationally efficient way. The caustic patterns enable compressed sensing, which exploits sparsity in the sample to solve for more 3D voxels than pixels on the 2D sensor. Our 3D voxel grid is chosen to match the experimentally measured two-point optical resolution across the field-of-view, resulting in 100 million voxels being reconstructed from a single 1.3 megapixel image. However, the effective resolution varies significantly with scene content. Because this effect is common to a wide range of computational cameras, we provide new theory for analyzing resolution in such systems.
1. INTRODUCTION
DiffuserCam is a compact, inexpensive lensless camera for single-shot volumetric imaging that uses compressed sensing to exceed conventional single-shot sampling limits. The prototype reconstructs large 3D volumes from one 2D exposure while characterizing scene-dependent resolution.
- Scanning and multi-shot systems offer high spatial resolution but sacrifice speed and hardware simplicity, whereas existing single-shot methods are fast but low-resolution.
- DiffuserCam places a diffuser before a sensor to encode volumetric 3D intensity into a single 2D image for sparsity-constrained computational reconstruction.
- 100 million non-uniformly spaced voxels are reconstructed from a single 1.3 megapixel image using a commodity-hardware prototype.
- The system provides true depth sectioning and 3D renderings while using simple calibration, no precise construction alignment, and light-efficient optics.
- Nonlinear reconstruction produces object-dependent performance, so standard two-point resolution can be misleading; a local condition-number analysis agrees with experiments.
- The compact device is intended to support high-resolution lensless 3D imaging of large and dynamic samples, with potential applications in remote diagnostics, mobile photography, and in vivo microscopy.
A. System Overview
DiffuserCam uses position-dependent caustic PSFs to encode 3D scenes into 2D measurements. A sparsity-constrained inverse problem recovers more voxels than sensor pixels when the object and optical encoding satisfy the model assumptions.
- A thin phase diffuser produces high-frequency pseudorandom caustic PSFs whose 3D position dependence captures volumetric information.
- Lateral source shifts translate the PSF, while axial shifts approximately scale it, giving each 3D position a unique pattern.
- Assuming mutually incoherent scene points, the sensor measurement is modeled as a linear combination of PSFs from sampled 3D positions.
- The forward model has sensor-pixel rows and a grid-determined number of columns, so full-resolution reconstruction can become underdetermined when voxels outnumber pixels.
- Compressed sensing imposes sparsity, nonnegativity, and optionally identity or total-variation transforms to recover the 3D object.
- Distributed, uncorrelated caustic patterns enable recovery because shifts and magnifications generate decorrelated columns in the forward matrix.
A. System Architecture
The prototype combines an off-the-shelf diffuser with a traditional sensor, using diffuser geometry and high-f-number optics to generate suitable caustics over a broad propagation range. Sensor binning produces 1.3 megapixel inputs for reconstruction.
- The hardware consists of a diffuser fixed a short distance in front of a sensor, with diffuser bumps acting like randomly spaced microlenses.
- The diffuser’s f-number sets the minimum caustic feature size and optical resolution, while its average focal length determines the highest-contrast caustic plane.
- The prototype uses a PCO.edge 5.5 Color camera and a Luminit 0.5° engineered diffuser with approximately 8 mm average focal length and f-number 50.
- High f-number preserves caustic contrast across propagation distances, so the diffuser need not be positioned precisely at the caustic plane.
- 2x2 binning matches sensor sampling to the caustic features and yields 1.3 megapixel images before 3D reconstruction.
B. Convolutional Forward Model
DiffuserCam models the 2D sensor measurement as a sum of depth-dependent point-spread functions, then exploits convolutional structure for scalable calibration and reconstruction. This formulation makes the underdetermined 3D inverse problem computationally tractable.
- Each voxel’s radiant power is mapped to a unique sensor caustic, and the measurement is the sum of all voxel contributions.The forward model represents the object on a non-Cartesian 3D grid and uses PSFs indexed by source position.
- A paraxial shift-invariance approximation converts lateral source translations into scaled lateral PSF shifts.For fixed depth, a source displacement (∆x, ∆y) produces a sensor displacement (m∆x, m∆y).
- The cropped convolution model avoids explicitly storing the enormous transmission matrix and enables 3D FFT-based operator evaluation.Without this reduction, the matrix would require petabytes of memory; the model instead evaluates equivalent convolutions efficiently.
- Calibration requires one on-axis caustic image per depth plane rather than measurements from every voxel.A single calibration image at each depth replaces millions of voxel-specific calibration images.
- ADMM and Fourier-domain diagonalization make large-scale inversions practical, although reconstructions still require substantial computation and memory.A 537-million-voxel reconstruction takes 26 minutes and 85 GB of RAM on a 144-core workstation.
3. SYSTEM ANALYSIS
System analysis determines the camera’s field of view and sampling limits from geometry, sensor angular response, and depth-dependent resolution. The analysis also shows that scene complexity can reduce effective resolving power beyond two-point measurements.
- 3. SYSTEM ANALYSIS: The proposed local condition number analysis addresses why standard two-point resolution can misrepresent performance for complex objects.The authors analyze resolution, field of view, and convolution-model validity to choose a reconstruction grid for real-world objects.
- Resolution: The effective resolution depends on object complexity: sixteen points fail at the two-point spacing but succeed when separation increases.This object-dependent behavior motivates using more realistic resolution criteria than isolated point pairs.
- Field-of-View: The field of view is limited by either geometric deflection constraints or the sensor’s angular acceptance.The sensor cutoff αc is defined where response falls to 20% of its on-axis value.
- Field-of-View: Axial resolution degrades with depth until the hyperfocal plane, beyond which depth information is unavailable.Objects beyond that plane can still be reconstructed as 2D images.
- Field-of-View: The prototype’s axial field of view extends from 7.3 mm to the 2.3 m hyperfocal plane, with angular field of view of ±42° in x and ±30.5° in y.These angular limits use αc = 41.5° in x, αc = 30° in y, and diffuser deflection β = 0.5°.
B.1. Two-point resolution
Two-point experiments define best-case lateral and axial distinguishability, while the physical model extrapolates these measurements across the volume to set a non-uniform reconstruction grid.
- B.1. Two-point resolution: Two-point resolution is measured by reconstructing pairs of point sources at varying lateral and axial separations.Sources are synthesized from separately captured 1 µm pinhole images at 532 nm.
- B.1. Two-point resolution: A pair is considered distinguishable when the reconstruction contains at least a 20% intensity dip between the sources.The regularizer is disabled and full 5 MP sensor data are used to estimate best-case resolution.
- B.1. Two-point resolution: The system has highly non-isotropic resolution, with lateral resolution varying by depth and axial resolution determined by PSF support-width differences.The shift-invariance model allows localized measurements to predict two-point distinguishability throughout the volume.
- B.1. Two-point resolution: Voxel spacing is selected to Nyquist-sample the modeled 3D two-point resolution across the field of view.Voxel sizes vary with depth, densest near the camera; axial resolution reaches its limit near 2.3 m.
- B.1. Two-point resolution: Objects within 5 cm of the camera can be reconstructed with somewhat isotropic resolution.This near-camera range is where the authors place objects in practice.
B.2. Multi-point resolution
DiffuserCam’s effective resolution depends on scene complexity: resolving two points does not guarantee resolving a larger group at the same spacing. Increasing source separation restores distinguishability, with usable lateral resolution degrading by approximately 1.7× in the tested complex scene.
- A 4×4 grid of 16 point sources was tested at two lateral and axial separations.The smaller spacing matched the measured two-point resolution limit, while the larger spacing was ∆x=75µm and ∆z=448µm.
- At the two-point limit, the system could separate two points but not all 16 sources.The tested two-point resolution was ∆x=45µm and ∆z=336µm.
- Increasing source separation to ∆x=75µm and ∆z=448µm made all 16 points distinguishable.
- The usable lateral resolution degraded by approximately 1.7× as scene complexity increased.The experiment motivates a theoretical framework because standard resolution metrics cannot be applied blindly to computational cameras.
C. Local condition number theory
The paper introduces local condition number analysis to explain how object complexity affects effective reconstruction resolution. The analysis links ill-conditioning to noise sensitivity and predicts degradation that approaches a limit as the number of sources increases.
- The theory analyzes the forward model numerically to determine how well it can be inverted as object complexity changes.
- Local condition numbers quantify noise sensitivity in sub-matrices associated with known source locations.Ill-conditioned matrices make reconstructions more noise-sensitive and require more iterations to converge.
- Conditioning worsens as neighboring sources move closer together, making small amounts of noise more likely to prevent source resolution.
- As the number of sources increases, the local condition number approaches a limiting case rather than worsening without bound.This supports estimating complex-object resolution from distinguishability measurements using a limited number of point sources.
- The local condition number theory explains experimentally observed resolution loss and may apply to other computational cameras.
D. Validity of the Convolution Model
The convolution model approximates spatially varying point-spread functions as shift invariant across the field of view. Experiments support this approximation near the optical axis, while edge resolution declines unless exhaustive calibration is used.
- The convolution model holds relatively well across the field of view, with inner products greater than 75%.Similarity is particularly good within ±15° of the optical axis.
- Registered point-spread functions at 0°, 15°, and 30° were compared using inner products and normalized spot size.
- Exhaustive calibration would improve edge resolution but increase calibration and computational complexity.The convolution approximation trades some field-edge resolution for simpler calibration and more efficient computation.
- On-axis resolution is retained through ±15° but gradually degrades beyond that range.The spot size was estimated from the peak width of the cross-correlation between on-axis and off-axis PSFs.
4. EXPERIMENTAL RESULTS
Experiments demonstrate single-shot 3D reconstructions of a resolution target and a plant, while showing that effective resolution depends on object complexity and differs from two-point predictions.
- The system reconstructs 3D images on a 2048×2048×128 grid, with the usable region constrained by the angular field of view.The grid is sampled to approximately match measured two-point resolution, using 1.3 megapixel measurements after sensor binning.
- At z = 24 mm, the tilted resolution target resolves features 79 µm apart, substantially worse than the 50 µm two-point resolution but near the 75 µm 16-point resolution.Group 2 element 4 is easily resolved, while element 5 is barely resolved.
- The experiments support multi-point distinguishability as a more informative measure for complex objects than the standard two-point resolution criterion.The authors characterize resolution variation with object complexity and relate the observed degradation to their 16-point analysis.
- A small plant is reconstructed and rendered from multiple angles, demonstrating recovery of three-dimensional leaf structure.The plant reconstruction was cropped to 480×320×128 voxels.