Source-linked AI summary
Adaptive foveated single-pixel imaging with dynamic super-sampling
David B. Phillips, Ming-Jie Sun, Jonathan M. Taylor, Matthew P. Edgar, Stephen M. Barnett, Graham G. Gibson, Miles J. Padgett
TL;DR
Single-pixel imaging faces low frame-rates and motion-related noise, while changing pixel boundaries complicates motion estimation. The paper uses adaptive foveated sampling and dynamic super-sampling to capture fast-changing detail quickly while accumulating resolution in slower regions, including a four-fold reduction in imaging time for moving features.
Problem
Single-pixel techniques face low frame-rates and motion-related noise compared with conventional multi-pixel image sensors.
Method
The system uses spatially varying cells and changing pixel geometries so successive frames provide complementary spatial information, with a moving high-resolution fovea across the field of view.
Results
A factor-of-4 reduction in imaging time reduces conventional motion blur and pattern-multiplexing noise, while composite reconstructions can reach 128×128 high-resolution pixels at 0.5 Hz.
Takeaways & Limitations
The tiered super-sampling framework supports video streams whose resolution and effective exposure time adapt spatially to scene evolution and can extend to other sequential-correlation computational imagers.
Takeaways & Limitations
The linear-constraint reconstruction scales as O(N^3) and was performed in post-processing, while future performance depends on more sophisticated sampling algorithms.
Abstract
from arXiv · showhide
As an alternative to conventional multi-pixel cameras, single-pixel cameras enable images to be recorded using a single detector that measures the correlations between the scene and a set of patterns. However, to fully sample a scene in this way requires at least the same number of correlation measurements as there are pixels in the reconstructed image. Therefore single-pixel imaging systems typically exhibit low frame-rates. To mitigate this, a range of compressive sensing techniques have been developed which rely on a priori knowledge of the scene to reconstruct images from an under-sampled set of measurements. In this work we take a different approach and adopt a strategy inspired by the foveated vision systems found in the animal kingdom - a framework that exploits the spatio-temporal redundancy present in many dynamic scenes. In our single-pixel imaging system a high-resolution foveal region follows motion within the scene, but unlike a simple zoom, every frame delivers new spatial information from across the entire field-of-view. Using this approach we demonstrate a four-fold reduction in the time taken to record the detail of rapidly evolving features, whilst simultaneously accumulating detail of more slowly evolving regions over several consecutive frames. This tiered super-sampling technique enables the reconstruction of video streams in which both the resolution and the effective exposure-time spatially vary and adapt dynamically in response to the evolution of the scene. The methods described here can complement existing compressive sensing approaches and may be applied to enhance a variety of computational imagers that rely on sequential correlation measurements.
FOVEATED SINGLE-PIXEL IMAGING
The system reconstructs spatially variant images from single-pixel correlation measurements by reformatting Hadamard patterns onto cells with differing sizes. This preserves the measurement count while increasing central linear resolution relative to uniform imaging.
- Measurement principle: Single-pixel imaging reconstructs scenes from correlations between the scene and sequential masking patterns.The system uses structured detection: a DMD masks the scene image and a photodiode records transmitted intensity for each binary pattern.
- Measurement principle: Hadamard masks provide a linearly independent, orthonormal basis for critically sampling an N-pixel image with N measurements.Each binary mask transmits light from approximately half of the image pixels, maximizing the photodiode signal.
- Uniform-resolution baseline: Doubling linear resolution requires four times as many measurements, creating a resolution–frame-rate trade-off in fully sampled single-pixel imaging.The uniform 32×32 system uses 1024 pixels and reconstructs frames at approximately 10 Hz with differential measurements.
- Reconstruction: The spatially variant reconstruction accounts for cell areas with a diagonal matrix, producing spatially varying frequency cut-off and signal-to-noise ratio.Each diagonal element encodes the area of the cell containing the corresponding high-resolution pixel.
- Experimental comparison: Using the same measurement resource as uniform imaging, the foveated reconstruction doubles linear resolution in the central region while reducing peripheral resolution.The cat’s face is enhanced in the foveal region, at the expense of lower-resolution surroundings.
SPATIALLY-VARIANT DIGITAL SUPER-SAMPLING
Digital super-sampling combines spatially varying cell geometries across co-registered sub-frames to enhance resolution unevenly across the field of view. Weighted averaging provides fast fusion, while linear constraints continue improving peripheral resolution as more sub-frames are combined.
- Super-sampling principle: Successive sub-frames sample complementary spatial information because their pixel boundaries are modified between frames.The measurements are inherently co-registered, enabling digital super-resolution for regions known to remain static.
- Super-sampling principle: Four overlapping sub-frames with half-cell translations double the linear resolution within the regular-grid fovea.Peripheral cells are instead randomly repositioned, producing non-uniform resolution enhancement outside the fovea.
- Fusion strategies: Weighted averaging fuses four recent foveal sub-frames quickly, whereas linear constraints solve a self-consistent system using all available measurements.The linear-constraint reconstruction is equivalent to deconvolving the weighted-average result with an appropriate spatially varying PSF.
- Fusion strategies: With linear constraints, peripheral resolution continues improving as more sub-frames are fused, while the fovea reaches maximum resolution after four sub-frames.Weighted averaging reaches the same foveal maximum after four sub-frames but mainly smooths the periphery when additional frames are included.
- Performance trade-offs: 8 Hz central detail and 2 Hz resolution-doubled images are delivered simultaneously, while 36 peripheral sub-frames can approach uniform high resolution.Uniformly imaging the full field at 128×128 hr-pixels would reduce the global frame-rate to 0.5 Hz.
- Performance trade-offs: Linear-constraint resolution gains come at the expense of reconstruction speed, with the method scaling as O(N^3) versus O(N) for weighted averaging.The linear-constraint reconstruction was performed in post processing, although graphics processors and efficient matrix manipulation could potentially enable real-time operation.
FOVEA GAZE CONTROL
The fovea is repositioned using feedback from recent measurements, allowing motion tracking and detail estimation while composite reconstruction accounts for local scene changes. Blip-frames support motion detection and effective-exposure mapping, while Haar wavelets guide sampling toward high-contrast detail.
- Adaptive guidance: Real-time feedback can reposition the fovea toward moving objects or regions anticipated to contain high levels of detail.This mimics saccadic movement by selecting the high-resolution region from previously measured visual cues.
- Motion tracking: Changing cell boundaries aids super-sampling but complicates motion comparison, so fixed-boundary blip-frames are interlaced with the variable-resolution sub-frames.Consecutive blip-frames are compared to detect motion, using 16×16 pixels at an interlacing frequency of 2 Hz.
- Motion tracking: Blip-frame interlacing reduces the average frame-rate by only ∼7%, whereas avoiding blip-frames would halve the super-sampling rate.The alternative uses pairs of sub-frames with identical pixel footprints for change detection.
- Motion tracking: Motion-guided fixation follows a moving detailed sign at 8 Hz and applies four-sub-frame super-sampling between blip-frames.Each sub-frame is recorded in 0.125 s, and the fovea location is updated at 2 Hz.
- Dynamic reconstruction: Difference-map stacks estimate when regions last changed and determine how many sub-frames contribute to each local reconstruction.This produces a dynamic video with spatially varying effective exposure-time and enhanced detail relative to uniform-resolution imaging.
- Detail estimation: A single-tier Haar wavelet transform guides the fovea toward regions with the highest unsampled edge contrast.The demonstrated trajectory samples most scene detail after 8 sub-frames, requiring 50% of the time needed for uniform high-resolution sampling.
DISCUSSION AND CONCLUSIONS
The demonstrated adaptive foveated system enhances single-pixel data gathering by allocating spatially variant sampling to scene dynamics. It can adapt the resolution–frame-rate trade-off, operate beyond visible wavelengths, and complement more sophisticated image-analysis methods.
- Adaptive foveated imaging: Every frame delivers new spatial information across the field-of-view while rapidly recording changing features and accumulating detail in slower regions.The strategy assumes that only some regions change from frame to frame.
- Relationship to compressive sensing: The approach does not require prior knowledge of a sparse image basis and may be combined with compressive sensing during sampling and reconstruction.The authors propose concentrating under-sampled measurements on important regions and using scene knowledge to improve accuracy and reduce noise.
- Wavelength flexibility: The system was demonstrated at visible wavelengths and in the SWIR range of 800-1800 nm through perspex opaque to visible light.The number of independently operating fovea can also be increased when required by the scene.
- Potential applications: The technique may enhance computational imagers based on sequential correlation measurements, including systems for scattering-media transmission and wavelengths lacking multi-pixel sensors.It provides a flexible way to adapt the resolution–frame-rate trade-off to the dynamics of the scene.
- Limitations and future directions: Future performance depends on the sophistication of the algorithms that control scene sampling, beyond the relatively simplistic motion-tracking algorithm demonstrated here.Motion-flow, pattern-recognition, and machine-learning methods are identified as possible enhancements.
- Limitations and future directions: Spatially variant sampling and reconstruction can be paired with advanced image-analysis techniques designed for varied real-world conditions.
1. Correlation measurements in the Hadamard
The system reformats Hadamard patterns onto non-uniform cells, reconstructs space-variant sub-frames, and fuses successive measurements to recover enhanced detail. This design trades peripheral resolution for faster acquisition while reducing motion-related noise and supporting composite images from changing scenes.
- Space-variant sampling: A binary transformation matrix maps N non-uniform cells into an M-pixel high-resolution grid, allowing Hadamard vectors to be stretched to the cell geometry.The matrix columns identify which high-resolution pixels belong to each cell, with cell areas represented in a diagonal matrix.
- Performance: By lowering resolution in static regions, the system reduces the time needed to image moving features at a target resolution by a factor of 4, reducing motion blur and pattern-multiplexing noise.The conventional uniform-resolution video uses the same measurement resource but cannot resolve lettering on the moving sign or features on the background target.
- Image fusion: Weighted image fusion uses inverse cell area to emphasize higher-resolution peripheral data while incorporating measurements from all sub-frames to suppress noise.Within the fovea, equal weighting is used; in peripheral regions, smaller cells receive greater weight because they provide higher local resolution.
- Image fusion: A linear-constraint fusion method solves for the high-resolution image by combining cell-sum constraints from multiple sub-frames, but noise can be amplified at the highest spatial frequencies.Weighted least squares can tune the relative importance of constraints when measurement noise is significant.