Source-linked AI summary
Robust learning from noisy, incomplete, high-dimensional experimental data via physically constrained symbolic regression
Patrick A. K. Reinbold, Logan M. Kageorge, Michael F. Schatz, Roman O. Grigoriev
TL;DR
Purely data-driven discovery has struggled with noisy, incomplete, high-dimensional experimental data and inaccessible variables. The paper combines physical constraints, weak formulations, and ensemble symbolic regression to discover a parsimonious fluid-flow PDE from velocity measurements. The resulting model is quantitatively accurate across weakly turbulent conditions and enables reconstruction of pressure and forcing fields.
Problem
Purely data-driven methods have been demonstrated mainly for simple, low-dimensional, low-noise systems, while real-world PDE discovery may involve unidentified or inaccessible variables.
Method
The approach combines physical principles for constructing candidate libraries, weak formulations to reduce noise sensitivity and remove latent-field dependence, and ensemble symbolic regression for parsimonious selection.
Results
η reached as low as 0.02, and a parsimonious model was consistently identified over 17.8 ≲ Re ≲ 36 while enabling pressure and forcing reconstruction.
Takeaways & Limitations
The discovered PDE is interpretable, comparable with first-principles models, and supports reconstruction of latent fields from velocity-only measurements.
Takeaways & Limitations
The method depends on domain knowledge and sufficiently varying data; it struggles when the flow becomes nearly stationary at lower Reynolds numbers.
Abstract
from arXiv · showhide
Machine learning offers an intriguing alternative to first-principles analysis for discovering new physics from experimental data. However, to date, purely data-driven methods have only proven successful in uncovering physical laws describing simple, low-dimensional systems with low levels of noise. Here we demonstrate that combining a data-driven methodology with some general physical principles enables discovery of a quantitatively accurate model of a non-equilibrium spatially-extended system from high-dimensional data that is both noisy and incomplete. We illustrate this using an experimental weakly turbulent fluid flow where only the velocity field is accessible. We also show that this hybrid approach allows reconstruction of the inaccessible variables -- the pressure and forcing field driving the flow.
A HYBRID APPROACH TO MODEL DISCOVERY
The paper introduces a hybrid approach to model discovery that combines three components for identifying governing models from experimental data.
- The hybrid approach combines physical principles, weak differential-equation formulations, and ensemble symbolic regression.These components are described as the approach’s three key ingredients.
Constructing the model library
The model library is constructed by combining general physical assumptions with fluid-specific knowledge about velocity, pressure, forcing, and symmetry.
- Constructing the model library: Causality, locality, and smoothness motivate Volterra-series terms built from velocity, latent fields, and their derivatives.For this fluid, the latent fields are pressure and body forcing because velocity evolution depends on internal and external stresses.
- Constructing the model library: The velocity model includes advection, diffusion, linear damping, nonlinear velocity, and vorticity-dependent terms before physical constraints remove incompatible terms.The candidate form is truncated to low order and later constrained using incompressibility.
- Constructing the model library: Euclidean symmetry constrains candidate terms to vector-valued forms with coefficients that are constant in space and time.Isotropy constrains functional form, while uniformity constrains the coefficients.
- Constructing the model library: Helmholtz decomposition shows that pressure and forcing must be included as additional fields when velocity-only terms cannot satisfy the governing equation.The decomposition relates the scalar and vector potentials to pressure and forcing.
Weak formulation of the model
The weak formulation addresses both inaccessible latent fields and noise amplification by integrating the model against spatiotemporal weight functions.
- Weak formulation of the model: The strong form is unsuitable because latent-field terms cannot be evaluated and differentiation amplifies measurement noise.Pressure cannot be recovered through a pressure-Poisson equation here because the forcing is neither known nor divergence-free.
- Weak formulation of the model: The weak formulation evaluates candidate terms over spatiotemporal domains using selected weight functions.The integration measure is dΩ = dx dy dt, with n = 0 corresponding to the time-derivative term.
- Weak formulation of the model: Stacking the integrated terms produces a linear system Qc = q0 for the unknown model coefficients.The coefficient vector contains c1 through cN, while Q is formed from the candidate-term vectors.
Ensemble symbolic regression
Ensemble symbolic regression identifies a sparse model by iteratively removing weak terms and testing the stability of the resulting functional form across data samplings.
- Ensemble symbolic regression: Iterative thresholding removes candidate terms whose weighted contribution falls below a chosen threshold, producing a parsimonious model.The over-determined coefficient system is repeatedly solved after deleting terms with small contributions.
- Ensemble symbolic regression: The residual alone cannot establish the correct functional form because redundant terms may leave the residual unchanged.For incompressible flow, a divergence-dependent term can alter the model without changing the residual.
- Ensemble symbolic regression: Ensemble regression across different data samplings quantifies robustness of the functional form and coefficient estimates.The ensemble uses different distributions of integration domains and can also incorporate different data sets.
RESULTS
Across 17.8 ≲ Re ≲ 36, ensemble symbolic regression identified a parsimonious velocity model matching first-principles structure, with accuracy depending on temporal variation and threshold choice. The identified model also enabled reconstruction of pressure and forcing fields, including a forcing profile nearly indistinguishable from direct measurements.
- Model selection: Thresholds 0.1 ≲ ε ≲ 0.3 balance robustness and accuracy; higher thresholds degrade fit, whereas lower thresholds make the model sampling-sensitive and overfit.The ensemble used 30 random spatiotemporal-domain distributions.
- Model identification: η reached 0.02 for a parsimonious model identified across 17.8 ≲Re ≲36, with advection, viscous flux, internal-stress, and external-stress terms matching Navier–Stokes structure.The identified equation is ∂tu = c1(u · ∇)u + c2∇2u + c3u −ρ−1∇p + ρ−1f.
- Model identification: The identified model form is identical to a first-principles derivation under assumptions including divergence-free horizontal velocity, while fitted parameters remain close but not identical to theory.The fitted c3 is about 25% higher than its theoretical value, consistent with the increase estimated to match the observed critical Reynolds number.
- Latent-field reconstruction: The forcing profile reconstructed from the measured flow is almost indistinguishable from the Lorentz-force profile computed from direct magnetic-field measurements.The reconstructed forcing corresponds to the depth-averaged Lorentz force across the electrolyte layer.
DISCUSSION
The hybrid approach discovers an interpretable, quantitatively accurate PDE from noisy, incomplete measurements while reconstructing latent fields. Its success depends on physical constraints and sufficiently time-varying data, with limitations from domain knowledge and experimental regime.
- The approach discovers a quantitatively accurate, interpretable PDE from noisy, incomplete measurements and permits reconstruction of latent fields.The model can be compared with first-principles models, which capture mechanisms qualitatively but not quantitatively accurately.
- Locality, causality, and spatial symmetries constrain candidate libraries before symbolic regression selects a parsimonious model.Domain knowledge also helps eliminate dependence on inaccessible pressure and forcing fields.
- The method identifies sparse models accurately at higher Re, where the velocity varies in time and both spatial coordinates.The approach experiences difficulties as the flow becomes nearly stationary at lower Re.
- The reconstruction accuracy decreases near steady flow because the chosen weight functions cannot provide solvable constraints when time dependence disappears.This limitation is attributed to latent variables and measurement error rather than symbolic regression itself.
- The framework extends beyond a single parabolic PDE to systems of elliptic, hyperbolic, higher-order PDEs, and ordinary differential equations.Alternative linear-algebra methods can be used when temporal-evolution terms are absent.
Experimental system and data collection
The experiment measures horizontal velocity in a shallow electrolyte–dielectric flow driven by magnetic Lorentz forces. Data span spatial grids and long temporal intervals, while magnetic-field measurements provide a reference for forcing reconstruction.
- The two fluid layers are each 0.3 cm thick in a 17.8 cm × 22.9 cm rectangular container, with temperature control limiting viscosity fluctuations.The setup is designed to make the electrolyte-layer flow close to two-dimensional, although bottom no-slip conditions retain vertical variation.
- The shallow electrolyte–dielectric flow is driven by Lorentz forces generated by current interacting with an array of permanent magnets.The magnetic field was measured across seven horizontal planes for validating reconstructed forcing.
- Particle image velocimetry measures the horizontal velocity field, which is approximately divergence-free for Re ≲50.The vertical flow component is negligibly small in this regime.
- Each velocity dataset samples both horizontal components on a uniform grid for at least 600 s at 1 s temporal resolution.The flow ranges from periodic behavior at low Re to aperiodic behavior with shorter autocorrelation times at higher Re.
Integration domains and weight functions
The method uses centered rectangular integration domains and smooth weight functions to average noisy measurements, transfer derivatives away from the velocity field, and eliminate latent-field contributions.
- Dataset description: Table I describes the datasets used for symbolic regression, including mean Reynolds number and whether reported times are periods or autocorrelation times.Times marked with an asterisk denote temporal periods; unmarked times denote autocorrelation times.
- Integration domains: Centered rectangular domains are distributed across the data, with spatial widths slightly below the flow size and temporal widths below the dataset duration.The choices avoid noisier side-wall regions and limit overlap so the resulting linear-system rows remain independent.
- Integration domains: Integration reduces noise through averaging, while domain sizes are selected to balance bulk-data quality against linear independence.The spatial domains are large in both directions, whereas temporal extent is restricted to limit overlap.
- Weight functions: Derivatives are transferred onto smooth weight functions by integration by parts whenever possible, reducing the direct amplification of PIV noise.The weights and their required spatial derivatives vanish at domain boundaries to remove boundary terms.
- Weight functions: For nonlinear terms that prevent complete derivative transfer, remaining velocity derivatives are computed in Fourier space with windowing and low-pass filtering.This fallback handles terms such as ω^2u while retaining noise-control procedures.
- Weight functions: Weight functions are constrained to remove inaccessible fields: temporal oddness eliminates time-independent forcing, while additional constraints eliminate pressure.The construction uses scalar fields and envelope functions to satisfy these conditions.
Reconstructing the pressure and forcing field
After identifying a parsimonious model, the method reconstructs pressure and horizontal forcing from the vector field s using Helmholtz decomposition, with filtering and temporal averaging addressing noise.
- Helmholtz reconstruction: Pressure and horizontal forcing are computed from the model-derived vector field s using Helmholtz decomposition.This reconstruction follows symbolic-regression model selection and targets the latent fields inaccessible in the measurements.
- Helmholtz reconstruction: The reconstruction uses the Fourier-space divergence-related component of s for the pressure/forcing decomposition.The displayed expression identifies the relevant k·s and Fourier-space structure used in the reconstruction.
- Noise control: The latent-field reconstruction omits the weak formulation, so Fourier-transformed model terms are low-pass-filtered before spatial derivatives are computed spectrally.Frequencies beyond |kx| > 2k0 or |ky| > 2k0 are removed, with the cutoff chosen empirically to balance relevant modes and noise.
- Accuracy: Forcing is less accurate than pressure on noisy data because its reconstruction requires one additional derivative.Because the forcing is stationary in this experiment, temporal averaging substantially improves its accuracy.