Source-linked AI summary
SINDy-PI: A Robust Algorithm for Parallel Implicit Sparse Identification of Nonlinear Dynamics
Kadierdan Kaheman, J. Nathan Kutz, Steven L. Brunton
TL;DR
The paper addresses the noise sensitivity of implicit methods for discovering nonlinear dynamics, especially rational-function systems. It develops SINDy-PI, which tests candidate terms with parallel and constrained sparse optimizations plus model-selection guidance. The method is demonstrated on challenging implicit ODEs and PDEs, including systems using limited and noisy data.
Problem
Existing implicit-SINDy extensions can represent rational nonlinearities but are extremely sensitive to noise, limiting identification of implicit dynamics.
Method
SINDy-PI tests candidate left-hand-side terms using parallel and constrained sparse optimizations, recasting implicit identification into non-implicit sparse regression with model-selection guidance.
Results
SINDy-PI identifies challenging implicit and rational ODE and PDE dynamics from limited and noisy data, with over 10^5 greater measurement-noise tolerance than implicit-SINDy in one example.
Takeaways & Limitations
The framework makes implicit dynamics and rational nonlinearities accessible in systems previously unattainable with SINDy, including complex ODEs and PDEs.
Takeaways & Limitations
Library design remains a scope boundary because rapidly growing libraries make sparse regression ill-conditioned and can reduce robustness.
Abstract
from arXiv · showhide
Accurately modeling the nonlinear dynamics of a system from measurement data is a challenging yet vital topic. The sparse identification of nonlinear dynamics (SINDy) algorithm is one approach to discover dynamical systems models from data. Although extensions have been developed to identify implicit dynamics, or dynamics described by rational functions, these extensions are extremely sensitive to noise. In this work, we develop SINDy-PI (parallel, implicit), a robust variant of the SINDy algorithm to identify implicit dynamics and rational nonlinearities. The SINDy-PI framework includes multiple optimization algorithms and a principled approach to model selection. We demonstrate the ability of this algorithm to learn implicit ordinary and partial differential equations and conservation laws from limited and noisy data. In particular, we show that the proposed approach is several orders of magnitude more noise robust than previous approaches, and may be used to identify a class of complex ODE and PDE dynamics that were previously unattainable with SINDy, including for the double pendulum dynamics and the Belousov Zhabotinsky (BZ) reaction.
1 Introduction
SINDy discovers interpretable sparse dynamical models, but its generalized linear formulation does not naturally represent implicit dynamics or rational nonlinearities. SINDy-PI addresses this gap with parallel candidate testing, convex optimization, and model-selection guidance.
- 1 Introduction: SINDy identifies parsimonious nonlinear models by selecting the fewest candidate terms needed to describe measured time evolution.Sparse models balance accuracy and efficiency while remaining interpretable by construction.
- 1 Introduction: Implicit-SINDy extends SINDy to rational dynamics, but null-space sparsity becomes extremely sensitive to even small measurement noise.Rational dynamics such as ˙x = N(x)/D(x) can be rewritten implicitly, although the trivial zero solution must be avoided.
- 1 Introduction: SINDy-PI recasts implicit-SINDy as a convex optimization problem by testing each library term as a possible known left-hand-side term.Removing the selected term converts the problem to a non-implicit sparse regression, and the procedure is highly parallelizable.
- 1 Introduction: The framework explicitly targets rational nonlinearities and complex implicit PDEs, while adding parallel, constrained, control-input, Hamiltonian, and model-selection capabilities.The authors describe these extensions as part of a broader framework for complex implicit dynamics.
2 Background
SINDy represents dynamics through sparse regression on a library of candidate functions evaluated from trajectory data. Implicit-SINDy enlarges that library to include states and derivatives, enabling implicit differential equations and rational-function dynamics but introducing noise-sensitive null-space computations.
- 2 Background: SINDy assumes that measured state trajectories admit a sparse representation in a prescribed library of candidate functions.The active terms are encoded by sparse coefficient vectors.
- 2 Background: Trajectory data and numerically computed derivatives are used to evaluate the candidate-function library before sparse regression determines active coefficients.Each library column is a candidate function evaluated over measurement snapshots.
- 2 Background: The generalized linear model can be solved with methods including STLSQ, LASSO, SR3, SSR, and Bayesian sparse regression.The library can also be augmented with partial derivatives for PDE identification and external forcing terms.
- 2.2 Implicit Sparse Identification of Nonlinear Dynamics: Implicit-SINDy generalizes the library to functions of both x and ˙x, allowing identification of implicit differential equations and rational-function dynamics.These dynamics include systems such as chemical reactions and metabolic networks with separated timescales.
- 2.2 Implicit Sparse Identification of Nonlinear Dynamics: Null-space computations make implicit-SINDy non-convex and highly ill-conditioned for noisy data.The method seeks sparse coefficient columns in the null space of Θ(X, ˙X).
3 SINDy-PI: Robust Parallel Identification of Implicit Dynamics
SINDy-PI robustly identifies implicit dynamics by testing each library term in parallel, reformulating the problem into sparse non-implicit regressions, and selecting accurate parsimonious models. It improves noise robustness and data efficiency for rational ODE and PDE identification.
- Constrained optimization formulation: SINDy-PI bypasses ill-conditioned null-space calculations by removing each candidate term and solving the remaining sparse regression problem.The resulting formulation can use established SINDy solvers and is considerably more robust to noise than implicit-SINDy.
- Model selection: Testing every candidate function yields sparse, accurate models for correct terms and dense, inaccurate models for incorrect terms.The procedure is highly parallelizable, and multiple candidate equations can be cross-referenced for consistent sparsity patterns.
- Model selection: AIC, BIC, and Pareto-front approaches provide alternative ways to balance model accuracy and parsimony during model selection.Candidate models can also be cross-referenced and used to refine the library by retaining consistently selected terms.
- Implicit and rational dynamics: SINDy-PI handles rational dynamics by restricting candidate functions to derivative-weighted library terms and identifying separate sparse models for each state derivative.Several accurate sparse models may be compared to validate the selected terms.
- Noise robustness: Over 10^5 measurement-noise levels were handled by SINDy-PI relative to implicit-SINDy while still recovering the correct Michaelis–Menten model.Structure error counts incorrectly added or deleted terms, averaged over 30 noise realizations and displayed as distributions in Figure 3.
- Data usage: About 12 times less data was required by SINDy-PI than implicit-SINDy to identify the hardest yeast-glycolysis equation, while library normalization improved learning rates.Figure 4 evaluates success rates across randomly sampled training-data percentages, averaged over 20 runs.
- Implicit PDE identification: For the modified KdV equation, SINDy-PI accurately identified the rational gain term 2g0/(1 + u) at large g0, whereas PDE-FIND could not represent it in its library.As g0 increased, PDE-FIND compensated by tuning other coefficients at small values and overfit the library at large values.
4 Advanced Examples
SINDy-PI is applied to challenging implicit and rational dynamics, including mechanical systems, a reaction PDE, and conserved quantities. It identifies correct models under low noise, while larger noise can cause misidentification, and physical-law discovery remains difficult because library construction is challenging.
- 4.1 Mounted Double Pendulum: SINDy-PI identifies the equations of motion for a mounted double pendulum from noisy trajectory measurements when the noise is low.The double pendulum contains rational nonlinearities that the original SINDy algorithm cannot identify.
- 4.1 Mounted Double Pendulum: For larger noise, SINDy-PI misidentifies the double-pendulum dynamics but retains short-term prediction ability.
- 4.2 Single Pendulum on a Cart: SINDy-PI correctly identifies the actuated single-pendulum-on-a-cart model up to a noise magnitude of 0.01.The experiment applies control to the cart and measures both cart and pendulum states.
- 4.3 Simplified Model of the Belousov–Zhabotinsky Reaction: SINDy-PI correctly identifies the simplified Belousov–Zhabotinsky reaction dynamics despite strong coupling and implicit behavior.These features make the reaction challenging for implicit-SINDy and PDE-FIND.
- 4.4 Extracting Physical Laws and Conserved Quantities: SINDy-PI can extract physical laws and conserved quantities, but library construction is difficult because higher-order derivatives and many states enlarge the library.Large libraries make sparse regression sensitive to noise, and the paper presents only one physical-law example.
5 Conclusions and Future Work
The paper presents SINDy-PI as a robust extension for implicit dynamics and rational nonlinearities, using parallel and constrained optimization with model-selection guidance. Across challenging ODE, actuated, PDE, and physical-law examples, it reports substantial noise robustness and reduced data requirements, while library design and model selection remain important future challenges.
- SINDy-PI combines parallel and constrained optimization to identify implicit dynamics and rational nonlinearities while incorporating external forcing and actuation.The framework is demonstrated on ODEs, actuated systems, PDEs, and conserved quantities.
- The examples demonstrate considerable noise robustness and reduced data requirements compared with implicit-SINDy.
- Library size can grow rapidly, making sparse regression ill-conditioned and motivating automatic, expert-guided library generation as future work.The paper also identifies model selection as a continuing concern involving accuracy, sparsity, and overfitting.
A.1 Performance Evaluation Criteria
The evaluation compares SINDy-PI and implicit-SINDy using prediction, structural, and parameter accuracy against a ground-truth model.
- Models are selected by lowest prediction error on test data and compared with ground truth using prediction, structural, and parameter accuracy.
A.2 Numerical Experiments
The numerical experiments compare SINDy-PI with implicit-SINDy across noisy simulations using repeated noise realizations, model selection, and multiple accuracy criteria. SINDy-PI outperforms implicit-SINDy in all tested cases and uses substantially less data for one challenging equation.
- The experiments use 2400 training and 600 testing initial conditions, with 23 Gaussian noise levels ranging from 10^-7 to 5 × 10^-1.
- Derivatives are computed using finite differences and TVRegDiff, with TVRegDiff results shown because it produces more accurate derivatives.TVRegDiff requires hyperparameter tuning and causes aliasing, so the time-series ends are trimmed.
- SINDy-PI and implicit-SINDy are trained across sparsity parameters, then models are selected using test data and evaluated by prediction, structure, and parameter errors.Performance is averaged over 30 noise realizations at each noise level.
- SINDy-PI outperforms implicit-SINDy in all tested cases despite sensitivity to training length, prediction steps, initial conditions, library choice, and derivative computation.
B Data Usage of SINDy-PI and implicit-SINDy
The paper compares how much data SINDy-PI and implicit-SINDy need and describes evaluation setups for nonlinear ODE and PDE identification. Data-length experiments use shuffled trajectory samples, repeated trials, and prediction-error-based model selection.
- Yeast glycolysis: Data-usage experiments train SINDy-PI and implicit-SINDy on random percentages of shuffled yeast-glycolysis trajectories without added noise.The simulations use 900 random initial conditions with magnitudes from 0 to 3, time step dt = 0.1, and time horizon T = 5.
- Yeast glycolysis: Each data-length condition is evaluated over 20 numerical trials using the percentage of runs that recover the correct equation structure.The comparison focuses on the hardest yeast-glycolysis state equation; other state equations require less data.
- Modified KdV: The modified KdV comparison uses spectral-method data with dt = 0.01, T = 20, spatial domain L = −25 to 25, and n = 128 spatial points.PDE-FIND and SINDy-PI are compared using their respective candidate libraries.
- Modified KdV: For modified KdV, 80% of the data are used for training and 20% for testing and model selection, with 100 sparsity values from 0.1 to 10.The final model minimizes normalized prediction error for u_t.
- Simulation settings: The section also references simulation parameters for the double pendulum, single pendulum on a cart, and simplified BZ reaction, plus identified parameters under different noise magnitudes.The double-pendulum setup includes masses, geometries, inertias, angles, gravity, and friction coefficients; the single-pendulum model assumes no damping.
F SINDy-PI Models for the Mounted Double Pendulum
SINDy-PI is applied to a mounted double pendulum modeled through Lagrangian and Euler–Lagrange equations, including friction and coupled angular dynamics. The method recovers accurate equations at low and moderate noise but fails at higher noise.
- System formulation: The mounted double pendulum is parameterized by masses, center-of-mass positions, arm lengths, inertias, angular coordinates, gravity, and friction coefficients.The damping coefficients are k1 = 7.2485 × 10^-4 and k2 = 1.6522 × 10^-4.
- System formulation: The model is derived from a Lagrangian with a Rayleigh dissipation term, producing coupled Euler–Lagrange equations for the two pendulum angles.The symbolic equations include angular accelerations, damping, gravitational terms, and nonlinear coupling.
- Undamped dynamics: With k1 = k2 = 0, the coupled equations can be combined and solved explicitly for ¨φ1 and ¨φ2.Substitution of the numerical parameters yields rational expressions with trigonometric and velocity-dependent terms.
- Noise robustness: At noise magnitude 0.005, SINDy-PI identifies an equation whose reported coefficients closely match the corresponding noise-free model.The identified expression retains the rational denominator and the principal nonlinear terms.
- Noise robustness: At noise magnitude 0.01, SINDy-PI again identifies the rational structure with coefficients close to the noise-free expression.The reported model preserves the denominator and the main trigonometric and velocity-dependent terms.
- Noise robustness: At noise magnitude 0.05, SINDy-PI incorrectly identifies the dynamics and introduces additional terms.The reported expression contains extra harmonics and velocity–trigonometric products.
G SINDy-PI Model for the Belousov-Zhabotinsky Reaction
SINDy-PI identifies a simplified Belousov–Zhabotinsky reaction PDE with diffusion, linear, coupling, quadratic, and cubic terms. The discovered model is reported in explicit coefficient form.
- Discovered PDE: SINDy-PI discovers the simplified BZ reaction PDE as x_τ = ∆x + 0.24667x + 0.33333s + 0.5z + 3.3333xs − 5.0xz + 2.1333x^2 − 3.3333x^3.The model combines a Laplacian with linear, cross-species, quadratic, and cubic reaction terms.
H Inability to Identify Rational Dynamics with SINDy
The original SINDy algorithm cannot represent rational Michaelis–Menten dynamics globally with its polynomial approximation, whereas SINDy-PI and Implicit-SINDy recover the correct model. SINDy’s approximation works only near the expansion point and degrades with noise.
- Model mismatch: For Michaelis–Menten dynamics, SINDy identifies a Taylor expansion rather than the rational model, correctly recovering only its first three terms.The higher-order Taylor terms have large parameter errors.
- Model comparison: SINDy-PI and Implicit-SINDy identify the correct Michaelis–Menten model with highly accurate parameters on the tested trajectory.Their models match the true solution in the comparison, while SINDy agrees only near the origin.
- Noise and data dependence: Adding Gaussian noise further degrades the SINDy model, and identifying the correct model requires more data.The comparison uses 330, 2200, and 4400 training points spanning 0 to 12, with the same amounts used for model selection.
I Robustness of SINDy and SINDy-PI
SINDy-PI’s noise robustness varies with data, model, and library choices. In tests varying K_m, the maximum noise it handled increased as K_m increased.
- Noise robustness depends on data length, initial conditions, model parameters, and polynomial-term order.The relative impact of these factors varies across systems.
- Training data that explores more phase space and produces a better-conditioned Θ matrix generally supports more robust model discovery.Exciting transients and using ensembles of initial conditions are cited strategies for improving robustness.
- Higher-order polynomial libraries increase the condition number of Θ and make nonlinear terms harder to disambiguate.Library size also scales exponentially with maximum polynomial order.
- As K_m increases, the maximum noise SINDy-PI can handle also increases.Tests used K_m values of 0.01, 0.1, 1, and 10 with matched training and testing procedures.
- The effect of K_m on SINDy-PI’s noise robustness is summarized in Table 8.