Source-linked AI summary

A Comprehensive Review of Digital Twin -- Part 1: Modeling and Twinning Enabling Technologies

Adam Thelen, Xiaoge Zhang, Olga Fink, Yan Lu, Sayan Ghosh, Byeng D. Youn, Michael D. Todd, Sankaran Mahadevan, Chao Hu, Zhen Hu

arXiv:2208.14197v2cs.CEcs.AIcs.LG

TL;DR

Digital-twin research needs a more comprehensive account of the modeling, twinning, UQ, and optimization methods underlying the technology. This paper reviews the literature, proposes a five-dimensional definition, and classifies enabling techniques by physical-to-virtual and virtual-to-physical data flow. It covers modeling and twinning methods in Part 1, while the series extends to UQ, optimization, and a battery digital-twin case study in Part 2.

  • Problem

    A comprehensive review of modeling, twinning, UQ, and optimization methods enabling digital twins is missing from existing literature.

  • Method

    The paper reviews over 230 digital-twin studies, proposes five dimensions, and classifies enabling technologies by their roles and data-flow direction.

  • Results

    The review comprehensively examines physical-system modeling, P2V enabling techniques, and V2P actions and benefits across the physical system life cycle.

  • Takeaways & Limitations

    The five-dimensional definition organizes interactions between physical and virtual systems through data exchange, modeling, actions, and optimization.

  • Takeaways & Limitations

    Accurate solid models can require excessive time to create for physics-based simulations such as finite element analysis or motion simulations.

Abstract

from arXiv · show

As an emerging technology in the era of Industry 4.0, digital twin is gaining unprecedented attention because of its promise to further optimize process design, quality control, health monitoring, decision and policy making, and more, by comprehensively modeling the physical world as a group of interconnected digital models. In a two-part series of papers, we examine the fundamental role of different modeling techniques, twinning enabling technologies, and uncertainty quantification and optimization methods commonly used in digital twins. This first paper presents a thorough literature review of digital twin trends across many disciplines currently pursuing this area of research. Then, digital twin modeling and twinning enabling technologies are further analyzed by classifying them into two main categories: physical-to-virtual, and virtual-to-physical, based on the direction in which data flows. Finally, this paper provides perspectives on the trajectory of digital twin technology over the next decade, and introduces a few emerging areas of research which will likely be of great use in future digital twin research. In part two of this review, the role of uncertainty quantification and optimization are discussed, a battery digital twin is demonstrated, and more perspectives on the future of digital twin are shared.

1 Introduction

Digital twins connect physical systems with individualized digital models through continuous information exchange, supporting unit-specific decisions and optimization. This review addresses gaps in prior reviews by examining modeling and twinning methods and organizing them by data-flow direction.

  • Motivation: Digital twins use individualized digital models to characterize unique products or process units and optimize decisions for each unit.The concept differs from decisions based on average population characteristics.
  • Motivation: Physical-to-virtual data from sensors and inspections and virtual-to-physical actionable information close the digitalization loop.The loop supports control and maintenance actions tailored to individual units.
  • Research gap: Existing reviews address definitions, enabling technologies, and domain applications, but a comprehensive review of modeling, twinning, UQ, and optimization methods remains missing.The paper positions this gap as relevant to effective industry-scale digital-twin implementation.
  • Review scope: Based on more than 230 research papers, the review categorizes modeling and twinning methods according to their roles in digital twins.It also discusses UQ, optimization, challenges, and future research directions across the two-part series.
  • Review scope: Part 1 covers digital-twin definitions, research trends, modeling technologies, physical-to-virtual twinning, and virtual-to-physical actions.Part 2 focuses on UQ and optimization and includes a battery digital-twin case study.

2 Definitions of digital twin and literature overview

The paper reviews evolving digital-twin definitions, proposes a five-dimensional model organized around physical and digital systems plus data-flow and optimization functions, and surveys research growth across application domains.

  • Definitions: Digital-twin definitions evolved from virtual representations synchronized with physical systems toward broader, domain-specific formulations.The literature includes aerospace-origin definitions and later efforts to consolidate numerous application-specific definitions.
  • Definitions: Ambiguity among digital models, digital shadows, and digital twins motivates clearer distinctions based on physical–virtual data flow.The paper notes that vague and inconsistent terminology may hinder industry adoption.
  • Five-dimensional model: The proposed five-dimensional model integrates a physical system, digital system, updating engine, prediction engine, and optimization dimension.The model extends earlier three-dimensional and five-dimensional formulations, with F(·) integrating all five dimensions.
  • Five-dimensional model: A battery-cell example shows sensor measurements updating a digital model that forecasts degradation and informs retirement decisions for the battery pack.The example connects physical-system sensing, physical-to-virtual updating, digital modeling, prediction, and decision support.
  • Literature overview: The review covers 230 non-predatory-journal papers, with manufacturing contributing 107 papers and leading the surveyed field.Manufacturing is followed by mechanical, civil, and structures domains in the reviewed distribution.
  • Literature overview: Outside manufacturing, aerospace, civil, and mechanical domains lead citation trends, while fault diagnostics and prognostics are prominent applications.Five of the 12 top-cited non-manufacturing papers focus on fault diagnostics and prognostics; others address optimization, management, and service.
  • Literature overview: Recent publication counts are much higher in manufacturing, while civil, mechanical, and energy research have grown rapidly after excluding manufacturing.The apparent flattening from 2021 to 2022 reflects that the review was written in early 2022.

3 Modeling enabling technologies

Digital twin modeling uses solid, geometry-scanning, visualization, and physics-based techniques to create digital replicas and support analysis, optimization, and interaction. These methods involve trade-offs among detail, computational cost, data flow, safety, and modeling effort.

  • Solid modeling: Solid modeling creates 3D representations of equipment and environments for production-line layout, sequencing, human-robot interaction, tolerancing, and quality-control applications.Examples include optimizing machining order and assembly-line layout, studying ergonomics, and simulating equipment deviations affecting finished products.
  • Solid modeling: Solid models add geometry-based optimization dimensions beyond sensor measurements but can require substantial time and detail to construct.Their reuse from original equipment manufacturers could reduce this burden and support long-term reliability, maintenance, safety, and upgrade activities.
  • Laser scanning: Laser scanning and LiDAR repeatedly measure surface distances and assemble the measurements into detailed 3D point clouds for geometry modeling.Applications include checking manufactured parts, mapping buildings and bridges, and supplying data for digital twin sensor networks.
  • Laser scanning: Laser scanning in a digital twin system significantly improved manufacturing performance while maintaining a safe working environment.Real-time data streaming between physical and virtual dimensions was associated with more productive manufacturing lines without compromising worker safety.
  • Physics-based modeling: Physics-based modeling includes finite-element, computational-fluid-dynamics, and multiphysics analyses for structural, thermal, fluid, and related digital-twin applications.Examples include vehicle suspension stress, cutting-tool wear, aircraft-wing dynamics, beam-bending measurements, and battery cooling.
  • Physics-based modeling: Choosing model fidelity requires balancing computational cost against input-output accuracy, because high-fidelity models are not always preferred.Higher-fidelity simulations can be more accurate but slower, while model bias from approximations may require offline quantification and online compensation.

3.3 Data-driven modeling

Data-driven modeling supports digital twins when physics is poorly understood or too computationally expensive, providing efficient representations for real-time prediction, state estimation, control, and maintenance decisions.

  • Motivation: Data-driven models address cases where physics is too complicated to model accurately or physics-based simulation is too expensive for repeated digital-twin runs.Repeated runs are especially important for uncertainty quantification tasks.
  • Model classes: Digital-twin data-driven models include degradation models, surrogate models, and system-identification models.Surrogate models typically use simulation data, whereas system identification primarily uses online sensor or offline experimental data.
  • Model categories: Statistical and machine-learning models support both prediction and inference, with statistical models emphasizing project-specific probability models and ML emphasizing prediction.The paper further organizes data-driven models by their numerical or analytical construction.
  • Statistical models: Statistical system-identification methods include linear approaches such as ARMA, ARIMA, and SSI, alongside nonlinear NARX and NARMAX models.NARX represents current output using past outputs and current or past exogenous inputs; NARMAX additionally includes past noise terms.
  • Digital-twin roles: Identified dynamic-system models can enable real-time state estimation and control, while statistical degradation models support end-of-life prediction and predictive maintenance.Physics-based degradation models may be unavailable when degradation mechanisms are complicated or incompletely understood, requiring empirical statistical models.
  • Limitations: Statistical degradation models face challenges integrating continuously collected sensor data and quantifying uncertainty for maintenance decision making.These challenges are particularly relevant for high-value physical assets and safety-critical applications.
  • Deep learning: Deep-learning autoencoders can learn low-dimensional state-space representations from high-dimensional inputs for nonlinear system identification.The encoder maps high-dimensional inputs to a low-dimensional state vector that can be more readily viewed and interpreted.

3.4 Physics-informed ML

Physics-informed ML combines first-principle knowledge with data-driven learning to improve physical-system modeling, but representative hybrid approaches remain incomplete and face adoption barriers.

  • Overview: Hybrid physics-ML approaches combine first-principle simulation data with physical-system data when both sources are available.The review broadly classifies these approaches according to how the first-principle model is integrated with the ML model.
  • Approach 1: Physics-Informed Loss Function: Physics-informed loss functions penalize ML predictions that violate first principles, constraining training toward physically compliant solutions.This is described as the most common hybrid physics-ML approach and does not necessarily require physics-simulation training data.
  • Synthetic-data approaches: Physics-based synthetic data can partially constrain ML outputs, improve generalization, and reduce training-data requirements for large-scale deep-learning models.Synthetic-data generation and related hybrid strategies are among the six approaches summarized in Fig. 10.
  • Approach 5: Delta Learning: Delta learning adds data-driven residuals to physics-based or physics-trained ML predictions to compensate for missing physics and unit-to-unit variability.The approach combines baseline physics for broad operating-condition generalization with residual corrections for individual systems.
  • Challenges: The review notes that its hybrid physics-ML list is incomplete and that practical adoption is impeded by limited realistic datasets and benchmark problems.These datasets are needed for training, validation, testing, and benchmarking generalization and interpretation.
  • Taxonomy: The six-approach taxonomy includes physics-informed loss, synthetic-data generation, pretraining and fine-tuning, first-principle correction, residual correction, and learned first-principle inputs.Fig. 10 marks reported improvements in generalization and interpretability with separate icons.

3.5 System modeling

System-modeling techniques represent interactions and interdependencies among digital-twin components, with UML, SysML, ontologies, and knowledge graphs providing complementary structural and semantic abstractions.

  • Overview: System modeling captures high-level interactions and interconnection structures among the many interacting subsystems and components of a physical system.The review focuses on widely used techniques for representing these interactions.
  • Unified Modeling Language (UML): UML provides a standardized graphical language for specifying software structure and visualizing interactions among models.Its use in digital twins emphasizes interaction modeling on the software side.
  • Systems Modeling Language (SysML): SysML extends and omits UML elements to provide discipline-independent architecture modeling across software, hardware, personnel, and facilities.Its four pillars are structure, behavior, properties, and requirements.
  • Systems Modeling Language (SysML): SysML supports design, specification, analysis, verification, and validation through cross-system extensibility and requirements traceability.Examples include tracking lifecycle design changes and fusing domain models in shop-floor digital twins.
  • Digital-twin architecture: UML and SysML provide a unified interface for managing software complexity and describing digital-twin module structure, behavior, and interactions.An architecture-level representation of interdependencies eases understanding of object-oriented digital-twin software systems.
  • Ontology and knowledge graphs: Ontologies define domain concepts, entities, attributes, and interrelationships, while ontology maps represent generalized entity types rather than specific instances.Knowledge graphs instantiate ontology concepts with specific data through nodes and edges representing entities and relationships.
  • Ontology and knowledge graphs: Ontology and knowledge-graph representations provide a common language for identifying physical assets and supporting digital-twin evolution during system or asset-management changes.An ontology-based framework was used to enable co-evolution between a digital twin and its physical asset.

4 Physical-to-virtual (P2V) twinning enabling technologies

Physical-to-virtual twinning methods update digital-twin models from physical-system information and are related to the modeling methods reviewed in the paper.

  • P2V overview: The review summarizes five categories of physical-to-virtual twinning methods.These methods vary according to the modeling approach used.
  • Measurements as input: Measurements are identified as an input category for physical-to-virtual twinning.The supplied section heading names measurements as an input but provides no further method-level detail.
  • P2V overview: Fig. 12 relates modeling methods to physical-to-virtual and virtual-to-physical twinning methods.Some P2V methods, such as probabilistic model updating, apply across multiple modeling approaches.

4.1 Physical measurements as input to the virtual space

Physical measurements provide a direct physical-to-virtual connection by updating digital models or supplying inputs to physics-based analyses and designs. These approaches support real-time prediction, risk assessment, decision making, and planning, but direct state updates apply only when monitored quantities correspond to digital states.

  • Physical measurements can serve as inputs that let virtual-space digital models predict physical-system responses in real time.This supports timely risk assessment and decision making through the virtual-to-physical direction.
  • One updating strategy uses measurements to update digital-model movements or positions, supporting behavior modeling, disturbance detection, and simulation-based planning.
  • A second strategy uses monitoring data as inputs for physics-based analysis or design.
  • Data-transfer methods for sensor-based physical-to-virtual connections must match application requirements, because digital-twin simulations may require high transfer speeds for real-time optimization and control.Connections are broadly categorized as wired or wireless.
  • Direct physical-to-virtual updating is limited to cases where monitoring data can directly update digital states; uncertain engineering states require more advanced twinning methods.

4.2 Probabilistic model updating

Probabilistic model updating represents digital states as evolving, partially observed quantities and recursively estimates them from noisy measurements. Bayesian filters and dynamic Bayesian networks provide flexible mechanisms for uncertainty-aware updating, with computational choices depending on model structure and assumptions.

  • A digital state is a time-varying set of variables characterizing the digital models, such as equipment health, structural parameters, or battery state.
  • Digital-twin state estimation distinguishes physical and digital states, often calibrating most digital parameters offline while updating only a small identifiable subset online.
  • A state-space model represents state evolution through a transition function and process noise, while measurements relate observations to the hidden state through a measurement function and measurement noise.The state vector x_k, inputs u_k, process noise ω_k, observations y_k, and measurement noise v_k are defined in this formulation.
  • Recursive Bayesian filtering estimates the current hidden state from current and past noisy measurements by approximating its conditional distribution through transition and update steps.
  • Alternative methods that approximate the full conditional distribution may be computationally costly as the time horizon grows and are unnecessary for many digital-twin applications.
  • Kalman filtering is exact for linear Gaussian state-space models, whereas the extended Kalman filter approximates nonlinear or non-Gaussian cases by linearization and Gaussian-noise assumptions.
  • Dynamic Bayesian networks generalize hidden Markov models by representing time-evolving variables and can fuse heterogeneous data through flexible probabilistic dependencies.Their conditional-independence structure expresses likelihoods through local conditional probability relationships, and they have been proposed as a unifying digital-twin formulation.

4.3 ML model updating

ML model updating addresses operational changes that can shift deployed data away from training distributions. The review describes drift detection, online updating, Bayesian parameter estimation, and physics-informed strategies for maintaining model generalization and reliability.

  • Operational changes such as sensor modifications, process updates, or measurement-device changes can make test data differ substantially from the training distribution.
  • Data drift changes the input distribution while preserving the relationship between inputs and targets, causing models trained on the original distribution to make erroneous test predictions.
  • Concept drift changes the underlying input–target relationship between training and test data.The review illustrates this change using a regression problem.
  • Drift creates robustness and reliability challenges, so digital twins monitor performance with application-specific indicators and update ML models when needed.
  • ML model updating can be formulated as state-space parameter estimation and solved online with Bayesian filters such as extended Kalman or particle filters.
  • Physics-informed ML is presented as another route to improve generalizability when the underlying physical processes are known.

4.4 Fault diagnostics and failure prognostics

Fault diagnostics identifies current damage states, whereas failure prognostics tracks their evolution and forecasts future outcomes. The review connects these tasks to digital-twin health management, contrasting data, physics, and hybrid approaches while emphasizing data scarcity, feature-engineering burden, and real-time constraints.

  • PHM seeks to manage component and system health by reducing operational and economic impacts of failures and controlling maintenance costs.
  • Fault diagnostics detects, localizes, identifies, and may assess the severity of current damage, while prognostics tracks fault evolution over time.
  • ML-based fault diagnostics uses classification or prediction models to identify health states, fault types, or fault severity.
  • Conventional ML successes depend on domain knowledge for handcrafted feature extraction and selection, with feature engineering accounting for more than 90% of model-building effort.
  • Deep learning enables end-to-end diagnostics by automatically extracting complex representations from large sensor datasets, but faulty-state data scarcity can cause imbalance and poor accuracy.
  • Physics-based, data-driven, and hybrid approaches constitute the main categories of failure prognostics.
  • Physics-based prognostics offer interpretability and generalizability but face calibration, degradation-understanding, identifiability, and computational-cost limitations.These constraints can make real-time decision making difficult.
  • A battery digital twin can update a cell model with state-of-health measurements, forecast capacity degradation and remaining useful life, and support selecting when to retire the cell.

4.5 Ontology-based reasoning

Ontology maps and knowledge graphs structure empirical knowledge for semantic reasoning in digital twins. Their applications support inference across heterogeneous industrial data and product-quality relationships.

  • Ontology representation: Ontology maps and knowledge graphs represent empirical knowledge in structured form for semantic reasoning.Ontology markup languages include Web Ontology Language, XML Schema, and RDFS.
  • Reasoning methods: Knowledge reasoning methods are categorized into logic-rules-based approaches and related rule, pattern, path, and ontology-based techniques.Examples include predicate logic, injected rules, recurring patterns, Web Ontology Language reasoning, and path-ranking algorithms.
  • Digital-twin applications: Industrial ontology models combine business rules with analytics-derived knowledge to detect detrimental machining incidents.The resulting knowledge base serves as an inference platform for incident detection.
  • Digital-twin applications: Knowledge graphs integrate asset, environment, and system ontologies to support asset configuration, resource planning, and component maintenance.This integration accommodates heterogeneous data sources and drives inference of smart solutions.

5 Virtual-to-physical (V2P) twinning enabling technologies

V2P twinning technologies use virtual models and data-driven methods to guide control and maintenance actions across a physical system’s life cycle. The section focuses on model predictive control and predictive maintenance.

  • V2P overview: V2P connections include system reconfiguration, process control, production planning, maintenance scheduling, and path planning.The review concentrates on model predictive control and predictive maintenance as classical V2P technologies.
  • Model predictive control: Model predictive control uses a process model to predict future behavior and optimize constrained control actions.Its structure includes measurements, constraints, sampling points, and an objective function.
  • Model predictive control: Iterative finite-horizon MPC acquires current measurements, computes a cost-minimizing strategy, and reduces output error relative to a reference.The horizon must capture the effect of control commands on the controlled variable.
  • Model predictive control: Machine-learning-based MPC uses recurrent neural-network models to improve prediction accuracy and control performance for nonlinear dynamic processes.Future MPC may learn digital models directly from data, including deep architectures and data-driven predictive models.
  • Predictive maintenance: Predictive maintenance uses sensor data and preprocessing in an ML pipeline to monitor machine health and support planned maintenance before failure.Relevant bearing-health signals include acoustics, vibration, oil-wear particle count, and temperature.
  • Predictive maintenance: Physics-based digital-twin models can generate synthetic faulty or run-to-failure data and estimate digital states and remaining useful life from sensor inputs.Hybrid ML assistance can support estimation of digital state variables affecting lifetime.
  • Predictive maintenance: A Baker Hughes system analyzes real-time pressure, vibration, and timing signals to classify pump health as normal operation, monitor closely, or maintenance needed.The classification supports maintenance planning, although predictive RUL capability is unclear.

6 Perspectives on modeling and twinning in digital twins

Digital-twin research must address limited and heterogeneous data, privacy constraints, domain shifts, and emerging learning-based decision methods. The reviewed perspectives emphasize shared learning, domain adaptation, and deeper reinforcement-learning integration.

  • Data availability and sharing: A single system often lacks sufficient condition-monitoring data because variable operating conditions and rare faults limit coverage.Data from multiple similar systems can improve representativeness, but stakeholders may be reluctant to share it.
  • Data availability and sharing: Federated learning aggregates local model parameters so a fleet model learns from distributed experience without directly sharing condition-monitoring data.Each system retains its data locally while updating and receiving a current model version.
  • Data availability and sharing: Federated learning can still expose sensitive information through model updates, and privacy-preserving approaches are needed.The passage also notes that federated learning has not yet been broadly applied.
  • Domain adaptation: Data-driven digital twins may lose performance under unseen operating conditions or across fleet units with different configurations and regimes.These differences create domain shifts and a synthetic-to-real gap, often without target labels for fine-tuning.
  • Domain adaptation: Domain adaptation maps source distributions to target distributions so algorithms can generalize across domains and learn domain-invariant features.Distribution alignment and domain discriminators are among the proposed approaches.
  • Domain adaptation: A further practical challenge is that source and target datasets may have different label spaces, especially for fault diagnostics.This limits direct transfer between systems with different experienced fault types.
  • Reinforcement learning: Deep reinforcement learning in digital twins is being explored for model updating, control, decision support, and policy learning.Future work includes integrating it into data-driven or hybrid-twin learning and strengthening decision-support and control connections.

7 Conclusion

The paper defines digital twins through five dimensions based on data flow and reviews the modeling, P2V, and V2P techniques that enable them. It positions UQ and optimization as the focus of the second paper.

  • Conclusion: The proposed five-dimensional definition comprises a physical system, digital system, updating engine, prediction engine, and optimization dimension.The dimensions describe interactions including data exchange, modeling, and actions.
  • Conclusion: The next paper addresses how uncertainty quantification and optimization can be incorporated into the three dimensions covered here.The two-part structure separates enabling-technique review from UQ and optimization analysis.
  • Conclusion: This paper reviews techniques for physical-system modeling, P2V updating, and V2P actions across the physical system life cycle.Examples of V2P benefits include model predictive control and predictive maintenance.
  • Conclusion: Accurate physical representation combined with continual P2V and V2P interaction forms a closed loop for digital-twin implementation.Part 2 examines robust performance through UQ and optimization techniques.

Appendix A: A generic particle filter algorithm

Particle filtering approximates posterior state distributions with weighted particles and recursively updates them through transition, weighting, and resampling steps.

  • Particle filters use sequential Monte Carlo to recursively perform Bayesian filtering for state estimation in state-space models and dynamic Bayesian networks.
  • The posterior is approximated as a weighted sum of Dirac delta functions centered at the particles.Each particle has an associated weight, and N_P denotes the total number of particles.
  • The algorithm transitions particles forward to form a prior, evaluates measurement likelihoods, normalizes weights, and resamples high-weight particles.Resampling replaces negligible-weight particles with copies of higher-weight particles to mitigate particle degeneracy.

Appendix B: Decomposition of likelihood and Bayesian inference in a DBN

Bayesian inference in a dynamic Bayesian network decomposes posterior state distributions through conditional probability relationships and recursively propagated priors. Particle filters are commonly paired with DBNs for digital-state updating because they impose fewer assumptions than Kalman filters.

  • The DBN posterior factors into conditional probability tables or distributions linking child states to their parent states, together with a prior distribution.The cited factors describe probabilistic causality between parent and child nodes.
  • The prior is obtained at each time step through recursive Bayesian inference and uncertainty propagation using observations and state-transition probabilities.
  • Particle filters are usually used with DBNs to update digital states because they require fewer assumptions about nonlinearities and noise distributions than Kalman filters.Surrogate models are also identified as part of the stated rationale for using particle filters with DBNs.

Authors’ contributions

The authors divided responsibility for the review’s technical topics among contributors, while all authors participated in manuscript writing, review, editing, and comment.

  • Hu, C. and Hu, Z. devised the original concept, while Hu, Z., Thelen, A., and Zhang, X. conducted the literature review.
  • Contributors were assigned specialized areas including geometric, physics-based, data-driven, physics-informed, and system modeling.
  • All authors read and approved the final manuscript and participated in writing, review, editing, and comment.
  • Additional responsibilities covered probabilistic and machine-learning model updating, diagnostics, prognostics, maintenance, MPC, federated learning, domain adaptation, and perspectives.
Loading 2208.14197v2…