Source-linked AI summary
Coded aperture compressive temporal imaging
Patrick Llull, Xuejun Liao, Xin Yuan, Jianbo Yang, David Kittle, Lawrence Carin, Guillermo Sapiro, David J. Brady
TL;DR
High-speed video capture is limited by electronic bandwidth and power, motivating compression that does not impose additional modulation overhead. The paper introduces CACTI, which mechanically translates a passive coded aperture for CDMA-based temporal compression and reconstructs high-speed frames from coded measurements. The experiments demonstrate reconstruction at NF = 148, while the implementation uses C = 14 because mechanical deceleration makes two mask positions inaccurate.
Problem
Electronic read-out limits high-speed imaging, while prior compressive approaches can increase operating bandwidth, system volume, power, or cost.
Method
CACTI mechanically translates a passive coded aperture to encode temporal channels, then uses CDMA signal separation and GAP iterative reconstruction to recover video.
Results
NF = 148 frames are reconstructed from coded snapshots with little aesthetic difference from NF = 14 in the reported experiments.
Takeaways & Limitations
The passive translated-aperture framework offers mechanical simplicity, large compression ratios, inexpensive scalability, and extensibility to other frameworks.
Takeaways & Limitations
The compression ratio is C = 14 rather than 16 because mechanical deceleration makes the triangle wave’s peak and trough inaccurate for the model.
Abstract
from arXiv · showhide
We use mechanical translation of a coded aperture for code division multiple access compression of video. We present experimental results for reconstruction at 148 frames per coded snapshot.
1. Introduction
CACTI addresses bandwidth and power limits in high-speed imaging by mechanically translating a passive coded aperture to compress space-time measurements without added code transmission. It recovers multiple high-speed video frames from one coded measurement through CDMA-based signal separation and iterative reconstruction.
- Motivation: Current electronics limit high-speed imaging because read-out power rises with pixel read-out rate, reaching 1-100W per sq. mm for full data-cube capture.Ambient illumination can support frame rates approaching 10^6 per second, but electronic read-out constrains practical operation.
- Motivation: Earlier compressive approaches reduce read-out bandwidth but can increase system volume, operating bandwidth, power, or cost.Spatial light modulators require substantial control signaling, while parallel camera arrays increase volume and cost by a factor of M.
- Proposed approach: CACTI mechanically translates a passive chrome-on-glass binary mask during exposure, avoiding code transmission and additional operating power for modulation.The mask is placed in an intermediate image plane and implements temporal coding through harmonic translation.
- Proposed approach: CACTI applies code division multiple access to separate temporal channels from compressed data and uses iterative reconstruction to estimate several high-speed frames from one coded measurement.The approach inverts a highly underdetermined system of equations to recover the temporal channels.
2. Theory
The coded aperture shifts across temporal channels, multiplexing high-speed space-time data into a low-speed snapshot that can be reconstructed by inverting an underdetermined linear system. Discrete mask motion and channel selection determine the reconstruction’s temporal sampling, approximation quality, and residual blur.
- Sensing model: Mechanical translation shifts a coded aperture across temporal channels before their integration into one detector image.The detected frame contains the summed, spatiotemporally multiplexed information from the coded channels.
- Continuous model: The coded aperture’s position s(t) during integration controls temporal modulation, while pixel sampling and integration impose spatial and temporal sampling factors.Without coding, the detector acts as a low-pass filter with resolution proportional to ∆x spatially and ∆t temporally.
- Continuous model: The moving code aliases higher video frequencies into the detector passband, potentially increasing the effective passband beyond the uncoded sampling limit.The increase depends on the spatial frequency support of the coded aperture transmission function.
- Sensing model: NF high-speed subframes are estimated from one snapshot g through a forward matrix H with N rows and N × NF columns.The matrix represents the three-dimensional transmission function as a two-dimensional linear transformation.
- Motion model: A discrete triangle-wave approximation models the mechanically translated mask; smaller d improves motion fidelity but increases the number of columns in H.Periodic triangular motion permits one forward matrix to reconstruct snapshots throughout the acquired video while respecting acceleration limits.
- Temporal channels: With d = 1, NF = C and each pixel integrates several nondegenerate coding patterns, whereas d < 1 gives NF > C and interpolates between critically encoded channels.The interpolated slices generally estimate motion direction but retain most residual motion blur.
3. Experimental Hardware
The prototype mechanically translates a passive coded aperture during camera integration to modulate space-time data, capture coded snapshots, and reconstruct high-speed video offline. Calibration samples mask positions to build the forward model, enabling comparisons up to 148 reconstructed frames per measurement.
- Hardware setup: The prototype combines a camera objective, chrome-on-quartz coded aperture on a piezoelectric stage, relay lens, and 640 × 480 monochrome CCD camera.The coded aperture is 5.06mm × 4.91mm and spans 248 × 256 detector pixels.
- Hardware setup: A 15Hz triangle wave moves the mask during integration, while a synchronized square wave triggers camera exposures during both motion directions.The frequency-doubled SYNC signal produces 30fps coded snapshots.
- Spatial and temporal modulation: The moving aperture applies local code structures across temporal channels, effectively shearing the coded space-time datacube and providing per-pixel flutter shuttering.Stationary mask images at positions d pixels apart simulate the motion in the forward matrix H.
- Forward model calibration: Calibration images the aperture under uniform illumination at discrete positions to account for system misalignments and relay-side aberrations when constructing H.The calibration uses zero-padding and mask positions spaced by d detector pixels.
- Forward model calibration: Using d = 0.99µm and storing every temporal channel produces NF = 160 reconstructed frames, while experiments compare NF = C = 14 and NF = 148.The forward model has dimensions 281^2 × (281^2 × NF) for both comparison cases.
- Experimental results: Up to 148 frames reconstructed from one exposure do not significantly reduce aesthetic quality or significantly affect residual error, while reconstruction time increases approximately linearly with NF.The piezoelectric stage is convenient for the prototype but is not considered the optimal translation mechanism; a low-resistance spring could use less power.
4. Reconstruction Algorithm
The reconstruction algorithm solves the underdetermined coded-measurement inversion with GAP, alternating data-fidelity and structural-sparsity projections. Its transform-domain sparsity model and convergence behavior support reconstruction as the number of recovered frames increases.
- Algorithm motivation: Increasing NF makes inversion difficult because the coded forward model multiplexes many local code patterns into a single discrete-time measurement.Least-squares and pseudoinverse methods cannot accurately reconstruct the resulting underdetermined systems.
- Algorithm motivation: GAP reconstructs the high-speed frames by exploiting structural sparsity in wavelet or DCT transform domains without requiring training data.The method is described as universal and insensitive to the data being inverted.
- Alternating projections: GAP alternates Euclidean projections enforcing data fidelity and structural sparsity, starting from θ^(0) = 0 until the estimate converges.The data-fidelity projection targets the linear manifold of frames whose coded integration matches the detector measurement.
- Alternating projections: GAP’s reconstruction improves monotonically over successive iterations under sufficient forward-model conditions, allowing users to stop for intermediate results and resume later.The algorithm is characterized as an anytime method.
- The Linear Manifold: The linear manifold contains solutions to the underdetermined measurement equations, which structural sparsity disambiguates.The manifold encodes legitimate high-speed frames integrated through the forward model to produce g.
- The Weighted ℓ2,1 Ball: The weighted ℓ2,1 ball is defined in transform-coefficient space, where an orthonormal transform rotates the sparsity constraint into voxel space.The transform may be a wavelet transformation or Discrete Cosine Transformation.
5. Results
CACTI reconstructs high-speed scenes from single coded snapshots, with experiments demonstrating temporal superresolution up to 148 frames and improved detail from moving-mask coding. Reconstruction quality remains visually similar across frame counts, but motion blur, ambiguities, and mechanical constraints limit performance.
- Reconstruction: Large temporal motion blur was reconstructed with GAP using DCT bases, whereas stationary reconstructions used wavelet bases.The moving code pattern can make features such as pouring water difficult to see directly.
- System constraints: The compression ratio is C = 14 rather than 16 because mechanical deceleration makes the triangle-wave endpoints inaccurate for the linear-motion model.The affected samples were excluded from H to reduce model error.
- Spatial resolution: Moving the mask during exposure captures more unique coding projections than a stationary mask, improving reconstruction quality for detailed scenes.A stationary binary aperture may block small features and produce difficult, artifact-ridden reconstructions.
- Experimental scenes: The reconstructed videos include eye motion, lens magnification, chopper-wheel lettering, and time-varying specularities in pouring water.The figure results compare NF = 14 and NF = 148 reconstructions.
6. Discussion and Conclusion
CACTI provides a mechanically translated coded-aperture framework for compressive high-speed video using conventional limited-bandwidth sensors. The paper presents GAP reconstruction, scalable passive coding, and possible extensions to higher-dimensional imaging, while identifying computational reconstruction as an ongoing challenge.
- Discussion and Conclusion: CACTI uniquely codes and decompresses high-speed video using conventional sensors with limited bandwidth.The framework is described as mechanically simple, scalable at low cost, and extensible to other frameworks.
- Discussion and Conclusion: GAP compactly represents sparse signals using selectable bases, requires no prior scene knowledge, and was used for all reconstructions except Figs. 12 and 13.The algorithm is described as fast-converging and scalable to larger images and compression ratios.
- Discussion and Conclusion: Large-scale CACTI implementations must reduce reconstructed data as well as transmitted data, and future work will adapt C to minimize computation for high-quality motion depiction.This identifies reconstruction cost as a remaining design consideration despite GAP’s computational efficiency.
- Discussion and Conclusion: Scaling CACTI to larger N requires a larger passive mask and detector sensing area, while translating the aperture provides C-fold temporal resolution without additional conventional-capture bandwidth.The comparison is with LCoS-driven strategies that modulate N pixels C times per integration.
- Discussion and Conclusion: Future systems may combine CACTI with higher-dimensional modalities such as spectral compressive video, including integration with CASSI for 4-dimensional datasets f(x,y,λ,t).This is presented as a future direction for large-scale imaging systems.