Source-linked AI summary

Nonlinear System Identification: A User-Oriented Roadmap

Johan Schoukens, Lennart Ljung

arXiv:1902.00683v1eess.SY

TL;DR

Nonlinear system identification lacks a single complete treatment because it spans broad system classes and user choices, while linear models can fail for nonlinear, time-varying behavior. This article offers an intuitive, example-supported roadmap centered on data, model selection, validation, experiment design, structural errors, and process noise. It concludes that nonlinear identification requires careful domain coverage and attention to structural errors and nonstandard noise effects.

  • Problem

    Linear system identification can become inadequate for nonlinear, time-varying systems, while nonlinear system identification spans too broad a field for a complete overview.

  • Method

    The article provides a user-oriented roadmap of nonlinear identification choices and challenges, illustrated with examples, guidelines, and supporting software.

  • Results

    The article presents nonlinear identification as an iterative process requiring experiment coverage, model-choice revision, and validation, with process noise and structural errors demanding specialized treatment.

  • Takeaways & Limitations

    Good nonlinear models can provide intuitive insight for design and control, but users must align experiments and model structures with the intended application and system behavior.

  • Takeaways & Limitations

    The topic is too broad for a complete overview, so the article’s topic selection and organization are strongly shaped by the authors’ personal journey.

Abstract

from arXiv · show

The goal of this article is twofold. Firstly, nonlinear system identification is introduced to a wide audience, guiding practicing engineers and newcomers in the field to a sound solution of their data driven modeling problems for nonlinear dynamic systems. In addition, the article also provides a broad perspective on the topic to researchers that are already familiar with the linear system identification theory, showing the similarities and differences between the linear and nonlinear problem. The reader will be referred to the existing literature for detailed mathematical explanations and formal proofs. Here the focus is on the basic philosophy, giving an intuitive understanding of the problems and the solutions, by making a guided tour along the wide range of user choices in nonlinear system identification. Guidelines will be given in addition to many examples, to reach that goal.

Summary

Nonlinear system identification addresses modeling problems where linear models cannot capture nonlinear, time-varying behavior. The article presents a user-oriented roadmap covering data, model choices, validation, experiment design, structural errors, and process-noise effects.

  • Motivation: Nonlinear system identification is needed when linear models become imprecise or fail to reproduce essential system behavior in nonlinear, time-varying applications.
  • Motivation: The article surveys applications across mechanical, electrical, telecommunication, automotive, biomedical, biological, and chemical systems.
  • Identification process: Identification depends on data, candidate models, estimation, and validation, with experiment design determining which system features inform the resulting approximation.
  • Model choices: Nonlinear model selection offers many structures, ranging from white-box physical models to black-box models, and is driven by user preferences and system behavior.
  • Nonlinear challenges: Complex nonlinear manifolds make experiment coverage, optimization initialization, and extrapolation avoidance especially important throughout identification.
  • Structural errors: Structural model errors can dominate noise disturbances, making classical linear-identification choices and results invalid and requiring appropriate corrective actions.

Goal of the (Nonlinear) System Identification Process

The identification process begins by choosing the modeling goal, model level, and application domain, then uses data and nonlinear analysis to guide those choices. The article emphasizes matching model complexity and structure to the intended use, available insight, and observed nonlinear behavior.

  • Modeling goal: Simulation models use inputs alone for new situations, whereas prediction models use measured outputs for one-step-ahead control applications.Simulation is harder because models can become unstable and require smaller structural errors than prediction models.
  • Modeling goal: Do not develop a complex simulation model when a simple prediction model meets the task, because good prediction may still fail to produce reliable simulation.
  • Model level: Select the model level by balancing physical or behavioral insight against the expense of building and using the model.Physical, semi-physical, and black-box choices span models built from physical descriptions to input-output mappings tuned directly from experimental data.
  • Application domain: Focus the model on signals and situations important to the application rather than attempting to cover every possible excitation.The article recommends defining a restricted application domain because an ideal model covering all applications is unattainable.
  • Data and nonlinear analysis: Periodic excitation can reveal, quantify, and separate even and odd nonlinear distortions while keeping experimental cost close to that of a linear study.For the forced Duffing oscillator, nonlinear distortions exceed the noise level by more than 40 dB, indicating potential improvement from a good nonlinear model.
  • Model structure: Block-oriented models separate linear dynamics from static nonlinearities, while parallel Wiener-Hammerstein structures can approximate fading-memory systems.Semi-physical models can also concatenate known submodels, and black-box predictors construct flexible input-output mappings directly from data.

Experiment Design

Experiment design must balance coverage of the intended operating domain, information about the selected model, and practical experimental cost. For nonlinear systems, optimal input design is difficult because the information matrix depends on higher-order input properties.

  • Domain coverage: The experiment should cover the full domain of interest because nonlinear models cannot reliably extrapolate beyond sampled regions.This coverage requirement may conflict with information-maximizing input design and is not guaranteed by excitation power-spectrum choices alone.
  • Practical design: Periodic excitations provide direct access to nonparametric noise and nonlinear-distortion analysis.Periodic signals also support validation across repeated, similar experiments.
  • Information content: The excitation amplitude distribution and power spectrum should maximize information about the parameters of the selected model structure.These choices must be balanced against persistency of excitation and other requirements for a well-behaved approximating model.
  • Multisine design: Multisine design sets the frequency resolution, period length, amplitude spectrum, and phases as user-controlled signal properties.Higher frequency resolution requires longer measurement time, while the amplitude spectrum should cover the frequency band of interest.
  • Design objective: Optimal experiment design seeks the minimum information needed for the modeling goal at the lowest experimental cost.Cost includes time, power consumption, and disturbance of the process.
  • Nonlinear optimal design: For nonlinear systems, the Fisher information matrix depends on higher-order moments and leads to a highly non-convex multivariate input-distribution design problem.The matrix may also depend explicitly on the unknown true parameters, which are often replaced by estimates from an initial experiment.

Model Validation

Model validation asks whether a model solves the intended problem on data it did not use for estimation. Nonlinear validation must examine domain coverage, structural errors, and the model’s behavior across relevant excitation classes.

  • Cross-validation: Cross-validation compares simulated or predicted model outputs with outputs from new validation data not used for estimation.Numerical fit measures quantify how much output variation the model reproduces.
  • Error diagnosis: An error level above the noise floor detects structural model errors that the user must assess for acceptability.Periodic validation excitations can indicate whether dominant errors arise from missed even or odd nonlinearities.
  • Higher-order validation: Higher-order whiteness and cross-correlation tests are required for full nonlinear validation but are often omitted because they are cumbersome and data-intensive.Higher-order moments become multidimensional objects, making visual evaluation across all lags difficult.
  • Excitation dependence: A nonlinear model tuned on tail data failed on sweep data because the sweep covered a much larger state domain despite similar input amplitude and power spectrum.The larger errors outside the tail-domain region illustrate the risk of extrapolation.
  • Uncertainty: Theoretically calculated uncertainty bounds can underestimate actual variability when structural model errors dominate noise disturbances.Repeated experiments with different excitation realizations are recommended to compare observed and theoretical standard deviations.
  • Domain coverage: Validation should use a rich data set covering the intended use and the full domain of interest.Phase-plane trajectories can help check whether estimation and validation experiments cover the relevant state domain.

Models

The examples show how physical insight and appropriate transformations can improve nonlinear models. Physical extensions capture overflow behavior in cascaded tanks, while resampling makes a buffer-vessel model linear in transformed data but nonlinear in the original variables.

  • Cascaded tanks: The cascaded tanks benchmark models water levels driven by a pump, with overflow producing hard nonlinearities and input-dependent process noise.The lower-tank output saturates at level 10 when overflow occurs.
  • Cascaded tanks: Overflow is modeled by constraining tank states to their maximum values and adding an extra process-noise term when the upper tank overflows.Additional loss terms provide flexibility that accommodates model imperfections beyond the expected fluid-loss effect.
  • Cascaded tanks: The extended physical model achieved fit = 1.02% on estimation data and fit = 1.78% on test data, versus 4.23% and 5.93% for the simple model.The extended model includes the additional loss terms.
  • Cascaded tanks: The Gaussian Process NARX reference achieved fit = 4.62% for simulation error and fit = 0.057% for one-step-ahead prediction error.The comparison illustrates that obtaining a small prediction error is easier than obtaining a small simulation error.
  • Buffer vessel: In the buffer vessel, changing volume and flow makes the system nonlinear, motivating a natural time variable based on Volume/Flow.The observed data are resampled according to this measured natural time variable.
  • Buffer vessel: The resampled-data linear model gives a much better fit than the initial linear model on the raw data.The semi-physical model is linear after resampling but nonlinear in the original raw data and sufficiently describes the buffer for time-marking pulp.

Examples III: Steel-Grey Model (Linearization-Based Model): The High

The HPFS example uses local linear model networks to represent nonlinear rail pressure dynamics, with experiment design focused on covering the operating regime. Compared with ramp-chirp excitation, OMNIPUS data produced better local agreement and highlighted the importance of informative excitation for generalization.

  • HPFS modeling: A local linear model network represents rail pressure as a nonlinear function of delayed engine speed, valve input, injection time, and pressure.The nonlinear function is represented through regime points, validity functions, and local linear models.
  • HPFS modeling: HILOMOT starts from a global affine model and adaptively splits the local model with the worst error into two submodels.Sigmoid validity functions are linked hierarchically to adjust spatial resolution.
  • Experiment design: The experiment must estimate local models while covering the full regime-point space under process constraints; OMNIPUS uses 10-minute, 100 Hz signals with 60000 samples.It is compared with a ramp-chirp excitation.
  • Results: The ramp-chirp LMN shows major mismatches in some operating regions, indicating that informative data are missing there, whereas the OMNIPUS LMN shows only a minor mismatch in the cited example.Both networks were evaluated on test data containing new ramp-chirp and OMNIPUS realizations.
  • Results: Gaussian-process and local-linear models behave similarly overall, but experiment design strongly affects generalization outside the training domain.In this example, the Gaussian-process models appear less sensitive to that issue.
  • Hydraulic crane example: 42% fit from the best linear crane model rises to 72% with an input nonlinearity, which is interpreted as a saturation-like hydraulic-to-force transformation.The linear model was judged insufficient for control design.

Examples V(a): Black Box Volterra Model of the Brain

The brain sensorimotor example identifies a regularized second-degree Volterra model relating wrist motion and EEG signals. The model captures nonlinear behavior while retaining substantial structural error, and richer excitation improves validation VAF without changing the qualitative conclusions.

  • Model and experiment: A regularized Volterra model represents the wrist-motion-to-EEG relation using multidimensional impulse responses constrained for smoothness and exponential decay.Regularization controls the rapidly increasing parameter count and permits identification from relatively short data sets.
  • Model selection: The maximum kernel degree is a critical choice: higher degrees rapidly increase parameters, whereas lower degrees create larger structural model errors.A prior analysis attributed more than 70% of the system characteristics to even nonlinear behavior, compared with 10% captured by a linear model.
  • Experiment and estimation: The experiment uses odd random-phase multisine perturbations, while EEG preprocessing removes frequencies below 1 Hz and 50 Hz mains disturbance before averaging.The averaged data have an SNR of about 20 dB and impose a 20 ms output delay.
  • Results: 46% average VAF is achieved on validation datasets across participants, while about 44% of output variance remains unmodeled after accounting for 10% noise.A fourth-order kernel could improve results, but richer and longer experiments were considered infeasible.
  • Results: The second-degree Volterra model still has large structural model errors but provides useful insights into the wrist-brain system.This makes the model informative despite incomplete predictive coverage.
  • Results: A richer excitation including even frequencies raises VAF to almost 60%, while the qualitative conclusions remain unchanged.With even frequencies excited, estimating the linear and nonlinear parts can no longer be decoupled.

Sidebar

The sidebar surveys ways to represent static nonlinear mappings, from breakpoint interpolation and basis expansions to neural networks, trees, Gaussian processes, and physically informed regressors. It also highlights practical trade-offs: known basis functions yield linear regression, while multivariate polynomial and Volterra expansions can create many parameters.

  • Static nonlinearities: Static nonlinearities map inputs to outputs at each time and serve as building blocks in state-space, regression, and block-oriented models.The same mapping can describe state transitions, regressors-to-output relations, or nonlinear blocks.
  • Breakpoint interpolation: Breakpoint parameterizations define function values at selected points and use interpolation rules such as piecewise constant or piecewise linear interpolation.More sophisticated interpolation rules are also possible.
  • Model complexity: Known basis functions make parameter estimation a linear regression, whereas multivariate polynomial and Volterra expansions can cause parameter counts to grow rapidly, reaching O(m^M) for Volterra models.This growth limits Volterra applications to shorter memory lengths and moderate polynomial orders unless regularization is used.
  • Basis expansions: Neural networks translate and dilate a mother activation function, with ReLU producing a piecewise linear and continuous mapping.Gaussian Bell and sigmoid functions are also listed as activation choices.
  • Trees: Regression trees partition the input axis through binary questions; leaves assign piecewise constant values, optionally followed by interpolation for piecewise linear functions.A depth-n tree has up to 2^n leaves.
  • Gaussian processes: Gaussian processes estimate a function through a posterior mean and quantify reliability using posterior variance, with smoothness encoded by the covariance function.Jointly Gaussian variables allow computation through linear algebraic expressions.

Sidebar

The sidebar presents periodic sine and multisine tests for detecting nonlinear distortions, alongside repeated-period statistics for separating nonlinear behavior from disturbing noise. It also distinguishes simulation from prediction models and explains when each is useful.

  • Nonlinear distortion tests: A sine test detects nonlinear behavior through harmonic components generated at the output by operations such as x^2 and x^3.The method is widely used but is described as not very robust.
  • Nonlinear distortion tests: A well-designed multisine excites selected odd frequencies so even nonlinearities appear at even frequencies and odd nonlinearities at unexcited odd frequencies.This robustifies and accelerates the sine test while making nonlinear distortions visible in the output spectrum.
  • Noise characterization: Repeated measurements of periodic input and output signals provide sample means and variances of disturbing noise by frequency, while periodic nonlinear distortions remain unchanged.Combining these quantities yields nonparametric information about the FRF, even and odd distortions, and noise spectrum.
  • Simulation and prediction: Simulation computes a model response from an input sequence, whereas prediction estimates future outputs from observed input-output data and can support control design.Prediction may use tentative future inputs for multistep forecasts.
  • Simulation and prediction: Prediction models use past outputs in addition to past inputs, while simulation models depend only on past inputs and can recursively replace measured outputs with predicted outputs.This relationship is illustrated in Figure S5.
  • Simulation and prediction: When disturbances are white noise, the simulation model provides optimal predictions; correlated disturbances can instead be partly predicted from their past.Whether an improved predictor is worthwhile depends on the nature of the disturbance.

Sidebar

Experiments on a forced Duffing oscillator illustrate how excitation choice, noise characterization, and model structure affect nonlinear identification. The examples show advantages of prediction, nonlinear models, and data-driven model reduction, while exposing complexity trade-offs.

  • Experimental setup: The forced Duffing oscillator is an electronic circuit emulating a nonlinear mechanical system with a cubic hardening spring and potentially rich dynamics.The setup is used throughout the article as an illustrative nonlinear benchmark.
  • Experiments: The arrow excitation has twice the maximum amplitude of the other signals and is used to test extrapolation, while all inputs share the same bandwidth.The tail and swept-sine signals have approximately the same RMS value.
  • Experiments: The sweeping-sine output becomes very large near resonance despite constant input level, creating internal extrapolation for models identified on the tail signal.This demonstrates that equal input levels do not imply equal operating conditions across experiments.
  • Noise and errors: 40 to 60 dB SNR corresponds to measurement and process noise at approximately the 1% to 0.1% level in the Duffing measurements.The input spectrum also contains spurious odd multiples of the 50 Hz mains frequency that act as process noise.
  • Linear models: Linear prediction can perform well under nonlinear distortions, whereas linear simulation is harder; prediction quality also depends strongly on the noise model.The better BJ noise model produces significantly smaller prediction errors than the ARX model.
  • Nonlinear models: For larger outputs, NARX errors are two to three times larger than NLSS errors, while the models are comparable at smaller output levels.The comparison is reported for the Duffing oscillator results shown in Figure S10.
  • Nonlinear models: Nonlinear prediction models outperform linear ones even without optimal tuning, and nonlinear models yield smaller simulation and prediction errors.The study attributes the latter comparison to better capture of system behavior by nonlinear models.
  • Model complexity: NARX parameter counts grow combinatorially with regressor number and polynomial degree, motivating pruning and structure-revealing methods.The cited example uses five regressors and degree three, and no further pruning was applied to those presented results.

Sidebar

The discretization discussion compares continuous- and discrete-time representations, emphasizing that simple finite-difference approximations can fail at relevant frequencies. Direct data-driven identification and tailored approximations provide more accurate alternatives for nonlinear systems.

  • Motivation: Discrete-time models are common in digital simulation and control, but careless approximation can produce behavior substantially different from the continuous-time system.The article first examines linear discretization before addressing nonlinear systems.
  • ZOH discretization: ZOH discretization exactly describes sampled input-output relations for linear systems, but this generalization is not always possible for nonlinear systems because ZOH behavior can be lost internally.The distinction motivates dedicated nonlinear discretization methods.
  • Finite differences: Finite-difference discretization works well only when the sampling frequency greatly exceeds the frequency band of interest; its relative error reaches 100% for f > 0.3f_s.The same rapid error growth is reported for the first-order-system approximation.
  • Direct identification: A discrete model fitted directly in the frequency band of interest has much smaller error than the finite-difference model for the continuous-time first-order system.The direct model is obtained by least-squares fitting over the specified band.
  • Nonlinear discretization: Theoretical analysis supports directly identifying explicit discrete-time models of continuous-time nonlinear systems, whose discretization error is dominated by aliasing effects.The direct-identification approach has smaller error than the TTS ZOH approach in the cited comparison.
  • Duffing illustration: In the Duffing study, errors decreased as O(1/f_s^4) until reaching the noise or structural-model-error floor, faster than 1/f_s.The drop rate followed the reduction in alias errors.
  • Practical choices: Data-driven discretization can require a more complex discrete-time model to keep approximation error below a user-defined level.Its approximation error is described as being dominated by aliasing.

Sidebar

Nonlinear models divide into NFIR systems with external dynamics and NIIR systems with internal nonlinear feedback. This choice affects representable behavior, stability analysis, structural detection, and identification complexity.

  • NFIR models use delayed inputs for external memory, whereas NIIR models use delayed outputs to create internal memory.
  • Unlike linear systems, NFIR and NIIR representations are not fully equivalent for nonlinear systems, so model choice affects structure, behavior, stability, and numerical methods.
  • Moving BLA poles under changing input offset or amplitude indicate NIIR structure, while fixed poles indicate NFIR structure.
  • NIIR models can represent chaos, bifurcations, jumps, moving resonances, hysteresis, and autonomous oscillations that NFIR models cannot.
  • Choosing NIIR rather than NFIR substantially complicates numerical optimization because parameter derivatives require evaluating another nonlinear model.
  • NFIR models suit fading-memory behavior, while strong nonlinear behaviors require internal-dynamics models such as NARX, feedback block-oriented, or nonlinear state-space models.

Sidebar

Structural model errors arise when the chosen model cannot represent the true system, and they depend on the applied input. They can invalidate classical uncertainty calculations and motivate identifying a sufficiently rich model before reduction.

  • Structural model errors remain even with infinitely many perfect data when the model structure cannot produce correct outputs.
  • Approximate models depend on the excitation class and are reliable only around the working domain represented by the applied inputs.
  • Experiments should cover the intended input domain, with power spectrum and amplitude distribution tuned so structural errors remain acceptably small.
  • In structural-error settings, classical covariance formulas underestimate estimate variability because structural errors depend on the input.
  • The simplified uncertainty analysis underestimates the observed standard deviation by 11 dB, about a factor of 3.
  • When affordable, identifying the most complex model first and then reducing it yields asymptotically efficient, consistent estimates and separates identification from application-oriented reduction.

– Linear Models of Nonlinear Systems

Linear approximations provide useful simplified descriptions of nonlinear systems, but their validity depends on the operating point or excitation. The BLA matches second-order behavior for a specified input and changes with input properties and process noise.

  • Linearization around an Equilibrium: Near an equilibrium, Taylor expansion yields a linear state-space approximation with transfer function G(s) = C(sI − A)^−1B.
  • Best Linear Approximation: The BLA is the linear model sharing the nonlinear system’s second-order signal properties for a specified input.
  • Best Linear Approximation: The BLA depends on the input spectrum, power spectrum, and amplitude distribution, so changing excitation conditions changes the approximation.
  • Experimental illustration: Four excitation levels in the Duffing oscillator show a resonance shift to higher frequency and noisier measurements caused by nonlinear distortions.
  • Process noise: Process noise entering before a nonlinearity systematically affects the output and the BLA relative to the known input.
  • Stochastic nonlinearities: For random excitation, stochastic nonlinearities are uncorrelated with but dependent on the input and can resemble measurement or process noise.
  • Process noise: Nonstationary periodic excitations can detect process noise; large detected process noise requires more advanced identification methods for consistent and efficient estimates.

Sidebar

The sidebar presents nonlinear state-space identification, stochastic inference methods, regularization, and decoupled representations as tools for managing model complexity and intractable likelihoods. These choices trade flexibility, interpretability, parameter freedom, and computational tractability.

  • Nonlinear state-space models: Nonlinear state-space identification estimates unknown parameters θ from observed inputs and outputs, with nonlinear dynamics, measurements, process noise, measurement noise, and latent states.The model uses f(·) for dynamics and h(·) for measurements, while θ denotes unknown static parameters.
  • Stochastic inference: Sequential Monte Carlo maintains weighted samples to approximate state distributions and can produce unbiased likelihood estimators for maximum likelihood identification.The weights are updated over time, and the estimate converges to the underlying distribution as the number of samples N approaches infinity.
  • Stochastic inference: Expectation maximization alternates approximation and maximization steps until convergence, reaching a stationary point while using smoothing distributions approximated by sequential Monte Carlo.This provides an alternative for solving maximum likelihood problems when process noise is present.
  • Complexity control: Black-box modeling offers flexibility with little physical insight but can create an exploding parameter count and increased overfitting risk as complexity grows.Regularization and data-driven structure retrieval are presented as ways to control parameter growth or its effect on modeled output.
  • Complexity control: Regularization can produce sparse or smooth solutions, but its success depends strongly on a regularization matrix P that encodes prior information or user preferences.Regularization restricts parameter freedom rather than necessarily reducing the parameter count, and introduces a bias–variance trade-off.
  • Decoupled representations: Decoupled representations reduce polynomial parameter growth from nqO(nd) to nqO(rd), improve interpretability through univariate branches, and may require fewer branches when nonlinear functions are tuned to the problem.Each branch is a univariate function gi(xi), with x = V ⊤p; the representation preserves a physical or intuitive interpretation while reducing parameters.

Sidebar Software Support

Nonlinear identification requires sophisticated software support because model definition, estimation, and analysis algorithms are typically complex. The sidebar surveys commercial and freely available packages, validation tools, analysis toolboxes, model-specific software, and benchmark datasets.

  • Software packages: The System Identification Toolbox commercially implements several methods and models discussed in the article.Its model objects include NARX, Hammerstein-Wiener, and user-programmed nonlinear state-space grey-box models.
  • Software packages: Validation and evaluation commands such as compare, resid, sim, predict, and forecast are available for nonlinear models and can also compare linear and nonlinear models.The commands support cross validation, residual analysis, simulation, prediction, and forecasting.
  • Free tools and datasets: FDIDENT supports nonparametric noise and distortion analysis, random-phase multisine design, and nonparametric nonlinear analysis.The article also identifies freely available software for polynomial nonlinear state-space models and documented experimental datasets through nonlinearbenchmark.org.
Loading 1902.00683v1…