Source-linked AI summary

On structural and practical identifiability

Franz-Georg Wieland, Adrian L. Hauber, Marcus Rosenblatt, Christian Tönsing, Jens Timmer

arXiv:2102.05100v2stat.MEphysics.data-anq-bio.QM

TL;DR

The paper examines identifiability in partially observed biological systems modeled with ordinary differential equations. It emphasizes models with well-determined parameters and predictions, while focusing biological insights on structurally and practically identifiable models.

  • Problem

    Partially observed biological systems often exhibit structural non-identifiability, while limited, noisy data can lead to bad models.

  • Method

    The paper translates biological systems into ordinary differential equations and proposes a more consequential use of methods for structural identifiability.

  • Results

    A good model has well-determined parameters and predictions, supporting biological insights from structurally and practically identifiable models.

  • Takeaways & Limitations

    Model assessment should prioritize the biological insights gained from models that are structurally and practically identifiable.

  • Takeaways & Limitations

    Because data is often recorded on a relative scale, models may require scaling and offset parameters for background correction.

Abstract

from arXiv · show

We discuss issues of structural and practical identifiability of partially observed differential equations which are often applied in systems biology. The development of mathematical methods to investigate structural non-identifiability has a long tradition. Computationally efficient methods to detect and cure it have been developed recently. Practical non-identifiability on the other hand has not been investigated at the same conceptually clear level. We argue that practical identifiability is more challenging than structural identifiability when it comes to modelling experimental data. We discuss that the classical approach based on the Fisher information matrix has severe shortcomings. As an alternative, we propose using the profile likelihood, which is a powerful approach to detect and resolve practical non-identifiability.

Introduction

Differential-equation models can turn qualitative biological reasoning into quantitative, predictive, and dynamic understanding, but over-parameterisation can leave parameters and predictions poorly determined. Identifiability and iterative model refinement distinguish useful models from models that merely fit data.

  • Differential-equation models can provide quantitative, predictive, and dynamic understanding of biological systems.
  • Useful models capture the main effects of the question, make experimentally falsifiable predictions, and enable biological insight.
  • Bad, good and useful models: Increasing model complexity until data can be fitted often produces over-parameterised models whose parameters and predictions are not well-determined.
  • Bad, good and useful models: Model refinement requires additional data and reduced or balanced complexity until the model has well-determined parameters and predictions.
  • The ultimate goal of systems-biology modelling is to use models to understand biology rather than to treat the model itself as the endpoint.

Parameter identifiability

Identifiability analysis is needed because limited, noisy measurements and partial observation can produce models with poorly determined parameters and predictions. The section distinguishes structural from practical identifiability and describes ODE-based observation and estimation.

  • Identifiability analysis supports constructing models with well-determined parameters and predictions, especially under limited and noisy biological measurements.
  • Structural identifiability concerns parameters indeterminable from model structure, whereas practical identifiability concerns insufficiently informative measurements.
  • Partially observed dynamical systems: Biological systems are represented by ordinary differential equations containing model states, unknown parameters, and external stimuli.
  • Partially observed dynamical systems: An observation function maps internal states to measured observations because not all components of biological systems can typically be measured.
  • Partially observed dynamical systems: Because the observation dimension m is typically smaller than the state dimension n, parameter estimation occurs in partially observed systems.
  • Partially observed dynamical systems: Parameter estimation commonly uses weighted residual sums of squares or negative log-likelihood under Gaussian errors, with maximum likelihood as a point estimate.

Structural identifiability

Structural identifiability asks whether model outputs determine a unique parameterisation. Non-identifiable parameters can vary while compensating changes preserve the output trajectory, creating parameter-space manifolds and links to non-observability.

  • A model is structurally identifiable when a unique parameterisation exists for every given model output.
  • Global structural identifiability requires the parameter to be uniquely determined across the entire parameter space.
  • A structurally non-identifiable parameter can change without altering the model trajectory because other parameters fully compensate for the change.
  • Local structural identifiability restricts the uniqueness assessment to a neighbourhood of the parameter value rather than the entire parameter space.
  • A model is structurally identifiable only when all of its parameters are structurally identifiable.
  • Structural non-identifiability implies a manifold of parameter values yielding unchanged outputs, while dynamic states may change along it, corresponding to non-observability.

A priori analysis of structural identifiability

Structural identifiability can be assessed with model-based a priori methods or data-based a posteriori methods, while profile likelihood characterises practical identifiability through parameter profiling and re-optimisation. The section also describes computational limits, confidence-interval behaviour, and Bayesian sampling difficulties.

  • A priori and a posteriori analysis: A priori methods assess structural identifiability from the model definition, whereas a posteriori methods use available data to find non-identifiable parameters.
  • A priori analysis: Many structural-identifiability methods rely on Lie groups, series expansions, differential algebra, differential geometry, or numerical algebraic geometry.
  • A priori analysis: Early structural-identifiability methods often apply only to low-dimensional systems because of computational complexity.
  • A priori analysis: A five-step pipeline combines numerical sensitivity analysis, symbolic calculations, re-parameterisation, model simplification, and verification of the re-parameterised model.
  • A priori analysis: In a model with 21 states and 75 parameters, two groups of non-identifiable parameters were detected and the model was re-parameterised.
  • A posteriori analysis: Profile likelihood varies one parameter around its maximum-likelihood estimate while re-optimising the remaining parameters, revealing coupled parameters.
  • Profile likelihood: Profile-likelihood computation can be demanding for larger systems because it requires numerical re-optimisation, motivating faster methods that avoid complete profiles.
  • Profile likelihood: An identifiable parameter has finite confidence bounds, a structurally non-identifiable parameter has infinite bounds, and a practically non-identifiable parameter can have one finite bound.

Re-parameterising structurally non-identifiable models

Recent computational advances make structural identifiability less of a bottleneck, although resolving many connected non-identifiable parameters remains difficult. Re-parameterisation can fix the problem but may sacrifice component-scale information and biological meaning.

  • Recent computationally efficient methods make determining structural identifiability less of a bottleneck for nonlinear ODE models.
  • Resolving many connected structurally non-identifiable parameters remains challenging, especially when few states are observed relative to the number of dynamic states.
  • Structural non-identifiability is usually addressed by re-parameterising the model, often by fixing some involved parameters.
  • Fixing parameters typically loses information about component scales, which can limit predictive power.
  • Biologically meaningful re-parameterisation after detecting non-identifiabilities remains challenging.

Practical identifiability

Structural identifiability guarantees practical identifiability only with infinite noiseless data, whereas practical identifiability concerns finite confidence intervals for parameter estimates and model predictions. The paper defines it through the combination of model and data.

  • Structural identifiability implies practical identifiability only with infinite data and zero noise.
  • Practical identifiability matters for obtaining precise parameter estimates and well-determined model predictions.
  • The literature has often treated practical identifiability vaguely, chiefly through the presence of large confidence intervals.
  • Here, a model and data are practically identifiable when confidence intervals for all estimated parameters are finite.

Parameter confidence intervals and identifiability

Profile likelihood provides confidence intervals suited to nonlinear ODE models and can reveal practical and structural non-identifiability. Fisher-information intervals can be inaccurate, uncontrollable, or insensitive to practical non-identifiability.

  • Profile likelihood provides a proper assessment of confidence intervals for estimated parameters.
  • Because nonlinear ODE solutions depend nonlinearly on parameters, applying Fisher-information confidence intervals to identifiability analysis is questionable.
  • Profile-likelihood confidence intervals can be asymmetric and invariant under model re-parameterisations such as logarithmic parameter transformations.
  • Adding enough measurements for existing observables can turn a practically non-identifiable parameter into an identifiable one.
  • Fisher-information intervals can be larger or smaller than profile-likelihood intervals and can become uncontrollable for finite measurements.
  • Fisher-information intervals can miss practical non-identifiability, falsely suggest structural non-identifiability, or be difficult to calculate because of singularity.

Bayesian methods for identifiability analysis

Bayesian sampling can agree with profile likelihood for structurally identifiable models but may mix poorly when structural non-identifiabilities exist and can be slower. Prediction profile likelihood instead propagates data uncertainty through prediction space.

  • Bayesian methods: MCMC-based practical-identifiability analysis is feasible only for structurally identifiable models because structural non-identifiabilities cause poor sampler mixing.
  • Bayesian methods: For structurally identifiable models, MCMC sampling yields results similar to profile-likelihood analysis.
  • Bayesian methods: MCMC was reported to be one order of magnitude slower than profile likelihood in a spatio-temporal reaction–diffusion model.
  • Bayesian methods: A comprehensive benchmark comparing MCMC and profile likelihood was missing.
  • Model predictions: Bootstrap and sensitivity-analysis approaches can require large numerical effort for nonlinear biological models with many parameters.
  • Model predictions: Prediction profile likelihood assesses prediction uncertainty by constraining the model response to a prediction z and exploring prediction space.
  • Model predictions: Practical identifiability can be increased by measuring additional data or reducing model complexity to match available information.

Achieving practical identifiability by new measurements with optimal experimental design

Practical identifiability can be improved by adding informative measurements, with optimal experimental design selecting targets and time points that constrain parameters or support model selection.

  • Practical identifiability can be achieved by adding new data.
  • Optimal experimental design searches for additional measurements containing maximal information about the system or selected parts.
  • For a specific parameter, trajectories along its parameter profile reveal measurement points with maximal information content, corresponding to high trajectory spread.
  • Prediction profile likelihood determines uncertainty at potential new measurement time points, promoting identifiability of the whole model.
  • Measurements at points with high prediction uncertainty effectively constrain the model, whereas low-uncertainty measurements better support model selection.

Achieving practical identifiability by reducing model complexity

When additional measurements are infeasible, reducing model complexity can address practical non-identifiability, although fixing parameters may reduce interpretative relevance. The review situates these strategies within an iterative modelling workflow aimed at biologically informative, identifiable predictions.

  • If additional data cannot be measured, model complexity must be reduced; parameters may be fixed using prior knowledge, sensitivity analysis, or profile likelihood.
  • Fixing parameters can decrease the interpretative relevance of model predictions, motivating systematic reduction strategies tailored to available data.
  • Likelihood profiles distinguish four non-identifiability scenarios based on profile flattening direction and whether other parameters are coupled.
  • Each scenario has a proposed cure: replacing a differential equation with an algebraic equation, lumping states, fixing a variable, or removing a reaction.
  • Model reduction conclusions should be documented for reproducibility, while unresolved directions include biologically plausible re-parameterisations and extensions to mixed effects models.
  • Identifiability analysis should use advanced methods, especially profile likelihood for practical identifiability, to assess model limitations and predictive power.
  • Practical non-identifiabilities can be detected by profile likelihood and addressed through model reduction or additional data.
  • Identifiability analysis is integral to an iterative modelling workflow whose goal is biological insight from structurally and practically identifiable models.

Conflict of interest statement

The authors declare no known competing financial interests or personal relationships that could have influenced the reported work.

  • The authors declare no known competing financial interests or personal relationships influencing the reported work.
Loading 2102.05100v2…