Source-linked AI summary
WIRE: Wavelet Implicit Neural Representations
Vishwanath Saragadam, Daniel LeJeune, Jasper Tan, Guha Balakrishnan, Ashok Veeraraghavan, Richard G. Baraniuk
TL;DR
Existing INRs can be accurate but are often slow, brittle to noise or parameter variation, and limited in fine-detail representation. WIRE addresses this by using a continuous complex Gabor wavelet activation, and experiments report state-of-the-art accuracy, training time, and robustness across diverse vision tasks.
Problem
Existing high-accuracy INRs remain slow, vulnerable to noise and insufficient measurements, sensitive to parameter choices, and limited in representing fine details.
Method
WIRE uses a continuous complex Gabor wavelet as the MLP activation, with complex weights and outputs preserving phase relationships while real signals use the real output part.
Results
WIRE achieves state-of-the-art INR accuracy, training time, and robustness across signal representation, inverse problems, and neural-radiance-field view synthesis.
Takeaways & Limitations
WIRE is presented as a go-to INR solution for signal representation and solving inverse problems because it combines favorable properties of prior nonlinearities.
Takeaways & Limitations
When complex weights are infeasible, WIRE can instead use the real or imaginary part of the complex Gabor wavelet.
Abstract
from arXiv · showhide
Implicit neural representations (INRs) have recently advanced numerous vision-related areas. INR performance depends strongly on the choice of the nonlinear activation function employed in its multilayer perceptron (MLP) network. A wide range of nonlinearities have been explored, but, unfortunately, current INRs designed to have high accuracy also suffer from poor robustness (to signal noise, parameter variation, etc.). Inspired by harmonic analysis, we develop a new, highly accurate and robust INR that does not exhibit this tradeoff. Wavelet Implicit neural REpresentation (WIRE) uses a continuous complex Gabor wavelet activation function that is well-known to be optimally concentrated in space-frequency and to have excellent biases for representing images. A wide range of experiments (image denoising, image inpainting, super-resolution, computed tomography reconstruction, image overfitting, and novel view synthesis with neural radiance fields) demonstrate that WIRE defines the new state of the art in INR accuracy, training time, and robustness.
1. Introduction
INRs offer continuous, flexible signal representations but existing methods face a tradeoff between accuracy, speed, robustness, and fine-detail representation. WIRE addresses this tradeoff with a complex Gabor wavelet activation and demonstrates strong performance across representation and inverse-vision tasks.
- Current INRs can take tens of seconds to fit high-dimensional data accurately, remain vulnerable to noise and insufficient measurements, and struggle with fine details.
- Harmonic analysis motivates wavelet atoms because typical vision signals are more concisely and robustly represented by atoms concentrated in space–frequency.
- WIRE replaces the MLP activation with a continuous complex Gabor wavelet, combining sine-like frequency compactness with Gaussian-like spatial compactness.
- Experiments report state-of-the-art INR accuracy, training time, and robustness across denoising, inpainting, super-resolution, CT reconstruction, image overfitting, and neural-radiance-field view synthesis.
2. Prior Work
Prior INR research spans CNN-based priors, continuous MLP representations, improved activations, and training accelerations. These approaches provide useful biases or speed, but high-capacity INRs can remain brittle, motivating WIRE’s Gabor-wavelet nonlinearity.
- Classical and data-driven regularization methods address noisy or ill-conditioned inverse problems, while CNN priors exploit architecture-specific image biases.
- INRs use continuous MLP-based function approximators, making them appealing for irregularly sampled signals such as point clouds.
- ReLU performs poorly for INR approximation, prompting positional encoding and sinusoidal or Gaussian nonlinearities, alongside Gabor-based multiplicative filter networks.
- Architectural changes such as adaptive decomposition, kilo-NeRF, and Laplacian-pyramid prediction accelerate INR training by leveraging multiscale visual structure.
- High-capacity INRs can train nearly instantly yet remain brittle because they overfit noise and signal alike; WIRE proposes complex Gabor wavelets to induce robustness.
- Wavelet transforms use translated and scaled short oscillating pulses, typically achieving faster approximation rates for signals and images than Fourier transforms.
3. Wavelet Implicit Representations
WIRE replaces standard INR activations with a continuous complex Gabor wavelet, combining frequency and spatial compactness in a wavelet-atom representation. Experiments and NTK analysis show improved accuracy, denoising convergence, parameter robustness, and multidimensional localization.
- 3.2. WIRE: WIRE uses the continuous complex Gabor wavelet ψ as its activation nonlinearity, with ω0 controlling frequency and s0 controlling spread.The first-layer activations are copies of the mother Gabor wavelet at scales and shifts determined by W1 and b1.
- 3.2. WIRE: WIRE represents signals with complex-valued weights and outputs, taking the real part for real signals while preserving phase relationships.Its complex exponential supplies periodic behavior, while the Gaussian factor supplies spatial compactness.
- 3.3. Implicit bias of WIRE: WIRE preferentially learns image signal over noise during early denoising optimization, converging orders of magnitude faster to essentially any given PSNR under NTK gradient flow.The analysis uses 64 × 64 × 3 Tiny ImageNet images with N(0, 0.052) i.i.d. pixel-wise additive noise.
- 3.3. Implicit bias of WIRE: WIRE converges an order of magnitude faster to the same PSNR than other INRs in ordinary denoising training.The evaluation uses 24 768 × 512 × 3 Kodak images with N(0, 0.052) additive noise.
- 3.4. Choosing the parameters ω0, s0: WIRE outperforms SIREN, Gaussian, and ReLU across a broad range of ω0 and s0 values for image representation and denoising.The reduced sensitivity to the exact parameters implies WIRE can be used without precise information about image or noise statistics.
- 3.5. Multidimensional localization: In two dimensions, WIRE’s first-layer activations resemble a mixture of Gabor wavelets and curvelets, providing more diverse spatial localization for natural images.This multidimensional localization significantly benefits the representation’s performance.
4. Experiments
WIRE achieves fast, accurate signal representation and performs robustly across denoising, super-resolution, CT reconstruction, and multi-dimensional inverse problems.
- Signal representation: WIRE reaches the highest representation accuracy for both images and occupancy volumes, attaining 43.2dB and 0.99, respectively.It also reaches high accuracy faster than competing approaches.
- Image denoising: WIRE produces the sharpest denoised image with the least residual noise from an input image at 17.6dB PSNR.Its qualitative result is similar to deep image prior.
- Super-resolution: WIRE reconstructs sharper single-image and multi-image 4× super-resolution results, preserving crisp and high-frequency details.For multi-image super-resolution, WIRE handles shifted and rotated observations on an irregular grid.
- CT reconstruction: WIRE yields the sharpest CT reconstruction from noisy, undersampled measurements, with pronounced features and fewer artifacts than competing nonlinearities.SIREN shows striation artifacts, whereas Gaussian produces overly smooth results.
- Multi-dimensional WIRE: 2D WIRE outperforms WIRE across denoising, CT reconstruction, and 4× super-resolution in PSNR and SSIM, while learning sharper features.The comparison uses equal parameter counts for WIRE and 2D WIRE.
5. Conclusions
The paper concludes that WIRE combines high representation capacity, fast training, and strong inductive biases for challenging inverse problems. Across the reported comparisons, its complex Gabor wavelet activation inherits favorable properties of several alternatives while providing a broad INR solution.
- Conclusions: WIRE combines higher representation capacity, faster accuracy gains, and strong inductive biases for challenging inverse problems.These advantages were validated through an extensive set of experiments.
- Conclusions: 2D WIRE achieves higher PSNR and SSIM than WIRE across the reported multi-dimensional inverse problems.The comparison covers denoising, CT reconstruction, and 4× super-resolution.
- Conclusions: WIRE is presented as a general INR choice because it combines favorable properties associated with SIREN, positional encoding, and Gaussian nonlinearities.The paper contrasts SIREN's capacity and speed, positional encoding's novel-view-synthesis suitability, and Gaussian's denoising favorability.
A.1. WIRE initialization
WIRE is largely robust to initialization, with only a marginal benefit from SIREN-like weights and up to 1 dB higher accuracy in the tested comparison.
- A.1. WIRE initialization: WIRE does not require specialized initialization and remains robust to its initial weights.It uses default uniform weights, unlike SIREN-like methods that strongly depend on initialization.
- A.1. WIRE initialization: SIREN-like initialization provides WIRE with up to 1 dB higher accuracy than standard initialization.The comparison covers image representation and image denoising with 20 dB input noise.
- A.1. WIRE initialization: WIRE’s initialization robustness enables easier tuning across a broad range of ω0 and s0 hyperparameters.
A.2. WIRE layer visualizations
WIRE produces sparse hidden-layer outputs for a Siemens star image, supporting accurate representation of high-frequency image regions.
- A.2. WIRE layer visualizations: WIRE uniquely produces sparse outputs in the second hidden layer compared with Gauss, SIREN, and ReLU with positional encoding.The Siemens star test image contains all spatial frequencies and orientations.
- A.2. WIRE layer visualizations: WIRE’s sparse outputs support higher approximation accuracy and sharper features in the image center’s highest-frequency region.Gauss is the next most sparse, while SIREN and ReLU with positional encoding produce blurrier center outputs.
A.3. Sensitivity to training parameters
WIRE maintains strong performance across learning rates, network depths, and feature widths, with particular advantages at practical small-to-medium model sizes.
- Effect of learning rate: WIRE achieves stable and significantly higher accuracy across a broad range of learning rates.Its highest tested accuracy occurs at a learning rate of 2 × 10^-2.
- Effect of number of layers: WIRE uniformly outperforms other nonlinearities as hidden-layer count varies, except with zero hidden features.For three or more hidden layers, its performance becomes similar to SIREN and Gaussian, while deeper networks are computationally expensive and often unstable.
- Effect of number of layers: WIRE is a reliable choice for small to medium numbers of hidden layers.This follows the reported tradeoff between its performance and the computational expense and instability of deeper networks.
- Effect of number of features: For more than 128 hidden features, WIRE outperforms other nonlinearities, with MFN a close second.At very low widths, all models similarly lack sufficient representational richness.
A.4. Inverse problems
WIRE performs strongly on inverse problems, achieving superior CT reconstruction with few projections and sharper multi-image super-resolution without ringing artifacts.
- Computed tomographic reconstruction: WIRE achieves higher PSNR than every tested nonlinearity for computed tomography reconstruction.Its reconstructions remain visually superior even with small numbers of projections.
- Computed tomographic reconstruction: WIRE’s visually superior CT reconstruction with few projections is beneficial for reducing X-ray exposure during capture.
- Multi-image super-resolution: WIRE generates the sharpest multi-image super-resolution features without ringing artifacts.It provides 1 dB or better reconstruction accuracy and 0.04 higher SSIM.
A.5. Neural radiance fields
WIRE provides robust and accurate neural radiance field reconstruction across training conditions, including varying learning rates, model sizes, projections, and numbers of training images.
- Effect of learning rate: WIRE remains robust to learning-rate variation and performs best with a high learning rate of 2 × 10−2 for noisy-image approximation.
- Effect of number of parameters: WIRE outperforms other nonlinearities with 128 or more hidden features and with one or more hidden layers.
- CT with varying number of projections: WIRE outperforms all other approaches by a considerable margin across computed tomography reconstructions with varying numbers of measurements.
- Effect of number of images: WIRE achieves the highest accuracy within 2500 epochs and converges more rapidly than other approaches on the drums dataset.
- Effect of number of images: 0.1dB higher than SIREN for 25, 50, and 75 training images, and 0.4dB higher with 100 images.
- Effect of number of images: WIRE generates sharper cymbal and stand features and a smoother drum membrane than the compared nonlinearities in novel-view reconstructions.