Source-linked AI summary
HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields
Lihao Chen, Xinyu Zhang, Panqi Chen, Lei Cheng, Ting Zhang, Jianlong Li, Shikai Fang
TL;DR
Sparse reconstruction of complex, oscillatory wave fields is severely underdetermined, especially with costly sensing and simulation. HarmoCore places a frequency-conditioned diffusion prior on compact Functional Tucker cores and performs posterior sampling in core space, achieving substantial gains across 2D and 3D benchmarks under 1%–2% sensing.
Problem
Reconstructing globally coherent, frequency-sensitive complex wave fields from scattered sensors is severely underdetermined, with sensing ratios as low as 1%–2%.
Method
HarmoCore jointly represents real and imaginary channels with compact Functional Tucker cores over shared continuous bases, then performs frequency-conditioned diffusion posterior sampling directly in core space.
Results
HarmoCore achieves the lowest Rel. L2 error and best physics consistency on all three benchmarks at 1%, 2%, and 5% sensing.
Takeaways & Limitations
The formulation is most effective in the most underdetermined regimes, where sparse observations leave the full field ambiguous and its margin over baselines is largest.
Takeaways & Limitations
The method relies on globally learned spatial parameterization and a sufficiently expressive Functional Tucker representation; equation guidance also depends on dataset-specific test-time metadata.
Abstract
from arXiv · showhide
Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem. Beyond the challenges of general physical-field reconstruction, wave responses are complex-valued, frequency-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme-sparse observations. Existing low-rank, operator, and diffusion approaches are largely designed for real-valued, smoother fields; dense pixel-space diffusion is particularly inefficient for oscillatory complex fields and difficult to scale to 3D. We propose HarmoCore, which places a generative prior in a compact, continuous, and structured wave-field latent. HarmoCore represents joint real--imaginary channels with Functional Tucker cores over shared continuous spatial bases, learns a frequency-conditioned core diffusion prior, and performs Diffusion Posterior Sampling directly in core space. At fixed sensor coordinates, the multilinear decoder induces an explicit likelihood guidance operator, avoiding dense pixel-space correction. Optional target-equation residual guidance further promotes physical consistency. Experiments on 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz show substantial gains under 1%--2% sensing while remaining practical in three dimensions.
1 Introduction
HarmoCore addresses the extreme underdetermination of reconstructing complex, frequency-sensitive, highly oscillatory wave fields from scattered sensors. It uses a structured functional latent diffusion prior and core-space posterior guidance to improve reconstruction under severe sparsity.
- Motivation: Existing approaches include low-rank fitting, sparse-to-dense regression, neural operators, and diffusion-based generative reconstruction, but most target real-valued, relatively smooth fields.The gap motivates a prior specifically aligned with oscillatory complex responses.
- Motivation: Wave-field reconstruction is difficult because real and imaginary components jointly encode amplitude and phase, while interference and frequency sensitivity produce rapid global variation.These properties make local interpolation unreliable when scattered sensors provide weak information about unobserved regions.
- Approach: HarmoCore represents each complex field with joint real–imaginary channels in a compact Functional Tucker core over shared continuous spatial bases.The bases encode continuous coordinate dependence, while the core retains sample- and frequency-specific coefficients.
- Approach: A frequency-conditioned diffusion prior is trained on the cores, and Diffusion Posterior Sampling uses a multilinear decoder-induced observation operator for sparse measurements.Optional equation-residual guidance can further promote physical consistency while correction remains in compact core space.
2 Preliminaries and Problem Setup
The paper formulates time-harmonic wave reconstruction as recovery of a continuous complex field from sparse channel measurements, then introduces Functional Tucker structure and diffusion posterior sampling for the inverse problem.
- Problem setup: A time-harmonic field is a steady-state complex spatial response whose real and imaginary parts jointly encode amplitude, phase, interference, node locations, and energy distribution.The target is the full continuous field, including unobserved regions.
- Problem setup: Each sensor provides a channel vector at a coordinate, while extreme-sparse experiments use sensing ratios as low as 1%–2%.Reconstruction quality is primarily assessed on the unobserved region.
- Problem setup: The fields satisfy governing equations involving a differential operator, medium parameters, and source terms, with metadata used for optional equation guidance.The Helmholtz operator is given as an example of the governing differential operator.
- Functional Tucker: Functional Tucker decomposition replaces discrete factor-matrix lookups with continuous coordinate-evaluable basis functions and a compact core.This reduces parameterization while separating shared spatial structure from sample-specific coefficients.
- Diffusion posterior sampling: Diffusion Posterior Sampling combines a learned diffusion prior with a measurement-consistency gradient evaluated on the denoiser’s clean estimate.Its efficiency depends on evaluating the measurement operator and gradient cheaply.
3 Method
HarmoCore learns a shared continuous-basis latent representation, models normalized joint-channel cores with a frequency-conditioned diffusion prior, and reconstructs cores through observation- and equation-guided sampling.
- Training: HarmoCore jointly learns shared continuous spatial bases and per-field compact cores from sparse observations, then trains a frequency-conditioned diffusion model on normalized cores.The two training stages separate structured representation learning from latent-prior learning.
- Latent representation: A single Functional Tucker core jointly stores real and imaginary channels, while shared sine-activated basis networks provide continuous decoding across samples, frequencies, channels, and coordinates.In 3D, a third basis network is added and the core gains a corresponding spatial mode.
- Latent representation: The training observation operator is formed by restricting the full-grid basis matrix to observed coordinates, so matrix multiplication evaluates decoded channels at sensors.A frequency-weighted spatial-smoothness regularizer is applied to the core matrices.
- Posterior reconstruction: At test time, frozen bases produce a reusable sensor-coordinate operator whose matrix–vector products provide the observation loss and core-space gradient.The same operator serves real and imaginary channels and is reused throughout reverse diffusion.
- Posterior reconstruction: Core-space reverse diffusion applies guidance to the clean estimate, combining observation consistency with an optional governing-equation residual correction.The equation term is supportive: removing observation-guided sampling causes a substantially larger accuracy collapse than removing the equation term alone.
4 Related Work
Related work spans deterministic sparse reconstruction, diffusion posterior sampling, and compact tensor representations. HarmoCore combines these ideas for extreme-sparse, frequency-sensitive wave fields using continuous functional cores.
- Sparse reconstruction: Sensor-to-dense networks and neural operators learn mappings from irregular or sparse observations, but their performance depends on training coverage of relevant response patterns.Under extreme sparsity, deterministic estimates can depend heavily on whether frequency and boundary configurations are represented in training.
- HarmoCore: HarmoCore applies Functional Tucker cores and diffusion to complex oscillatory wave fields, with the largest reported margins in the most underdetermined sensing regimes.The paper positions this combination as addressing frequency sensitivity and extreme sparsity together.
- Diffusion reconstruction: Diffusion models provide generative priors that can be combined with partial observations through posterior sampling and measurement-consistency guidance.This supplies a distribution of valid field configurations when local sensors do not resolve the global wave pattern.
- Functional representations: Tucker and tensor-train methods compactly decompose structured fields, while Functional Tucker models replace discrete factors with continuous-coordinate basis functions.Functional representations separate shared spatial variation from sample-specific coefficients.
5 Experiments
HarmoCore is evaluated across 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz under shared sparse-sensing protocols. It achieves its strongest advantages at 1%–2% sensing, with stable frequency behavior, useful physics guidance, and sensitivity to latent rank and distribution shift.
- Main Results across Benchmarks: HarmoCore achieves the lowest Rel. L2 error and best physics consistency across all three benchmarks at 1%, 2%, and 5% sensing.Its largest margin over the best baseline occurs at 1%–2% sensing, where the inverse problem is least constrained.
- Main Results across Benchmarks: At 2% sensing, HarmoCore’s Relative L2 Error stays low and comparatively flat across frequency, while baseline error fluctuates irregularly with ω.This indicates comparatively stable reconstruction behavior across the tested frequency range.
- Main Results across Benchmarks: HarmoCore is most valuable in the truly underdetermined regime, where its learned core-space prior resolves ambiguity left by extremely sparse observations.The margin over every baseline is largest when the sensor ratio is extremely low.
- Mechanism Analysis and Ablation: Removing DPS guidance causes a large accuracy collapse at both 1% and 2% sensing, whereas removing equation guidance causes a smaller but notable degradation, especially at 1%.The ablations identify observation-guided posterior sampling as primary and equation guidance as supplementary physical regularization.
- Mechanism Analysis and Ablation: A low PDE residual alone is insufficient for correct field recovery under extreme sparsity because equation consistency must be interpreted jointly with reconstruction error.The w/o DPS variant obtains lower PDE residual but substantially worse reconstruction accuracy.
- Additional Analyses: At 2% sensing, reconstruction error is non-monotonic in Functional Tucker rank, reaching its minimum at R=24 with 0.068 ± 0.039 and rising to 0.246 ± 0.106 at R=64.The rank sweep retrains the diffusion prior at each rank under the same architecture and training budget.
- Additional Analyses: On the 2D Synthetic out-of-distribution test set without retraining, HarmoCore’s error is essentially unchanged from its in-distribution values, while most other baselines degrade substantially.LRTFR is similarly robust, whereas DiffusionPDE shows the sharpest degradation among the remaining baselines.
6 Conclusion
HarmoCore reconstructs sparse complex wave fields using compact Functional Tucker cores and diffusion posterior sampling in core space. Across three benchmarks, it is most effective in highly underdetermined regimes, while its strongest evidence concerns sparse reconstruction rather than extrapolation or calibration.
- HarmoCore represents complex fields as compact Functional Tucker cores over shared continuous spatial bases and performs diffusion posterior sampling in core space.
- Across 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz, the formulation is most effective in the most underdetermined regimes.
- The method maintains better physical consistency than dense operator and pixel-space generative baselines.
- The method relies on globally learned spatial parameterization and a sufficiently expressive Functional Tucker representation.
- The paper provides stronger evidence for sparse reconstruction than for frequency extrapolation or uncertainty calibration, leaving those extensions for future work.
A Implementation Details
The appendix collects implementation details omitted from the main paper for space reasons.
- The appendix contains implementation details omitted from the main paper.
- These details were omitted because of space constraints in the main paper.
- The appendix serves as a supporting source for the paper’s implementation description.
A.1 Datasets and Preprocessing
The appendix describes scalar complex-field benchmarks in 2D and 3D, their sensing regimes, preprocessing, and the shared-basis latent diffusion implementation. Helmholtz data use PML-damped numerical solves, while the synthetic benchmark uses closed-form ray superpositions.
- Datasets and preprocessing: All three benchmarks use scalar complex fields with one real and one imaginary channel per sample-frequency pair.The experiments set K = 1 and C = 2.
- Datasets and preprocessing: The evaluation uses 153 held-out 2D Helmholtz fields, 170 held-out synthetic cases, and 410 held-out 3D Helmholtz cases.
- Datasets and preprocessing: Main-paper sensing ratios are 1%, 2%, and 5%, with 10% reported in the appendix.
- Datasets and preprocessing: Helmholtz benchmarks solve PML-damped equations on uniform 128^2 or 32^3 grids, while the synthetic benchmark uses direct-plus-reflected ray superpositions.
- Datasets and preprocessing: Helmholtz sources are random Gaussian point sources whose positions and phases remain fixed across each sample’s frequency sweep.
- Datasets and preprocessing: Shared SIREN spatial bases use default ranks Rx = Ry = 24, while a frequency-conditioned UNet diffuses over normalized joint-channel cores.The diffusion model uses scalar frequency conditioning through FiLM.
- Datasets and preprocessing: Richer frequency encodings did not yield consistent gains over scalar-FiLM conditioning.
- Datasets and preprocessing: At inference, frozen bases evaluated at sensor coordinates form a reusable observation operator for core-space guidance, with optional governing-equation residual guidance.
A.5 Evaluation Metrics
Evaluation uses full-grid relative L2 reconstruction error and physics residuals, with matched held-out data and specified baseline implementations. The appendix also documents the 10% sensing extension and synthetic benchmark’s absent PDE residual.
- Evaluation metrics: Relative L2 Error is the mean±std relative ℓ2 error over all channels and grid points, including sensor locations.
- Evaluation metrics: For Helmholtz benchmarks, the physics-residual metric evaluates the discretized governing operator against predicted fields and known sources on interior grid points.
- Evaluation metrics: All learned baselines use the same 80%/20% train-validation split and the same held-out test sets as HarmoCore.
- Baselines: FNO, F-FNO, VoronoiCNN, and DiffusionPDE are evaluated with the implementation settings specified for their respective architectures.
- Evaluation metrics: Table 4 reports reconstruction error and physics residual at 10% sensing, while the synthetic benchmark reports no governing-PDE residual.
- Baselines: LRTFR recovers per-case core coefficients by linear least squares from observed positions and decodes them on the full grid.
B.2 Out-of-Distribution Robustness (2D Synthetic)
The 2D Synthetic OOD test shifts wave speeds to a substantially wider range without retraining, while holding other generative factors fixed. Every evaluated baseline shows increased error under the shift, most sharply for DiffusionPDE.
- OOD test design: The OOD test redraws wave speeds from an empirically observed range of [0.44, 1.76], compared with [0.84, 1.19] in distribution.All other generative factors and the evaluation frequency grid are held fixed via the same random seed.
- OOD test design: No model is retrained; every method uses the same checkpoint evaluated in Table 1 on the shifted test set.Results are reported in Tables 3 and 5.
- Results: 0.038 → 0.140 at 5% sensing for DiffusionPDE, a more than 3× error increase under the distribution shift.VoronoiCNN, FNO, and F-FNO also show increased error, though less sharply.
B.3 Frequency-Conditioning Encoding for the Diffusion Prior: A Negative Result
The diffusion prior conditions on normalized frequency, with richer Fourier-style alternatives tested against the scalar default. On 2D Helmholtz at 10% sensing, the scalar-only encoding performs best, making the richer alternatives a negative result.
- Frequency conditioning: The default diffusion prior injects normalized frequency ωnorm ∈[0, 1] through FiLM at each ResBlock.A Fourier-style encoding with K = 8 frequency bands was also tested through the same conditioning path.
- Comparison: At 10% sensing, scalar-only conditioning attains the lowest reconstruction Rel. L2 error, 0.030, on 2D Helmholtz.Table 6 compares scalar, Fourier, polynomial, and linear spectral variants.
- Negative result: The richer encodings perform worse: Fourier-only reaches 0.049, while adding a linear spectral term to polynomial encoding reaches 0.052.Explicit polynomial terms and polynomial-only conditioning each reach 0.044.