Source-linked AI summary

A Comprehensive Review of Digital Twin -- Part 2: Roles of Uncertainty Quantification and Optimization, a Battery Digital Twin, and Perspectives

Adam Thelen, Xiaoge Zhang, Olga Fink, Yan Lu, Sayan Ghosh, Byeng D. Youn, Michael D. Todd, Sankaran Mahadevan, Chao Hu, Zhen Hu

arXiv:2208.12904v1cs.LGmath.OC

TL;DR

Digital twins require methods that connect physical systems with digital models while accounting for uncertainty and supporting optimization. This paper reviews enabling technologies, UQ, optimization, tools, datasets, challenges, and future directions, and demonstrates the approach with a battery digital twin for cell-retirement decisions. The case study combines degradation modeling and optimization to provide actionable information about a cell’s status across its life cycle.

  • Problem

    Digital-twin measurements and model outputs are uncertain, motivating methods to quantify uncertainty and support optimization across digital-twin dimensions.

  • Method

    The paper reviews UQ and optimization methods and constructs a battery digital twin integrating degradation modeling with optimization.

  • Results

    The battery digital twin provides actionable information for determining the optimal retirement time of a battery cell from its first-life application.

  • Takeaways & Limitations

    Combining degradation modeling and optimization can produce digital-twin information about a battery cell’s status across its overall life cycle.

  • Takeaways & Limitations

    Digital-twin validation is inherently lagging because validation data must come from the future use of the specific physical specimen being modeled.

Abstract

from arXiv · show

As an emerging technology in the era of Industry 4.0, digital twin is gaining unprecedented attention because of its promise to further optimize process design, quality control, health monitoring, decision and policy making, and more, by comprehensively modeling the physical world as a group of interconnected digital models. In a two-part series of papers, we examine the fundamental role of different modeling techniques, twinning enabling technologies, and uncertainty quantification and optimization methods commonly used in digital twins. This second paper presents a literature review of key enabling technologies of digital twins, with an emphasis on uncertainty quantification, optimization methods, open source datasets and tools, major findings, challenges, and future directions. Discussions focus on current methods of uncertainty quantification and optimization and how they are applied in different dimensions of a digital twin. Additionally, this paper presents a case study where a battery digital twin is constructed and tested to illustrate some of the modeling and twinning methods reviewed in this two-part review. Code and preprocessed data for generating all the results and figures presented in the case study are available on GitHub.

1 Introduction

This second paper reviews uncertainty quantification, optimization, and enabling technologies for digital twins, extending the modeling and twinning foundations established in Part 1. It also demonstrates these concepts through a battery digital twin and surveys industrial trends, open-source tools, datasets, challenges, and future directions.

  • Scope: The paper analyzes modeling and twinning technologies, uncertainty quantification, and optimization as complementary components of digital twins.It is the second paper in a two-part review series.
  • Relation to Part 1: Part 1 introduced a five-dimensional digital twin model and reviewed physical-to-virtual and virtual-to-physical technologies.The model is organized around data flow between physical and digital systems.
  • Roles of UQ and optimization: Uncertainty quantification and optimization support the synthesis of modeling, P2V, and V2P dimensions for functions including design optimization, quality control, and maintenance planning.These functions are considered in uncertain environments.
  • Case study: A battery digital twin demonstrates how reviewed methods can optimize the retirement of a battery cell from its first-life application.The case study illustrates predictive maintenance scheduling.
  • Perspectives: The paper closes by reviewing industrial-scale applications, open-source software and datasets, challenges, and future research directions.The review also discusses trends in industry.

2 Roles of UQ and optimization in digital twins

The paper organizes its discussion around uncertainty quantification and optimization methods used across digital-twin modeling and decision processes. It covers UQ for machine-learning and dynamic-system models, and optimization for sensing, physical-system modeling, and predictive decisions.

  • Uncertainty quantification: The review covers uncertainty quantification methods for machine-learning and dynamic-system models.It frames UQ as a core topic in digital twins.
  • Optimization: The review examines optimization methods for sensor placement, physical-system modeling, and predictive decision making.These applications span data collection, modeling, and decisions.
  • Integration: The section treats UQ and optimization as methods supporting digital-twin operation under uncertainty.The methods are discussed across multiple digital-twin activities.

2.1 UQ for digital twins

Uncertainty quantification addresses variability and incomplete knowledge across digital-twin dimensions, while supporting more reliable machine-learning and dynamic-system predictions. The review emphasizes epistemic uncertainty and argues that state-transition uncertainty can be especially important for extrapolation accuracy.

  • Uncertainty sources: Digital twins contain aleatory uncertainty from natural variability and epistemic uncertainty from limited data, knowledge gaps, and model simplifications.Aleatory uncertainty is generally irreducible, whereas epistemic uncertainty can decrease as information becomes available.
  • UQ of ML models: Machine-learning models may fail unexpectedly on out-of-distribution samples because their performance depends strongly on training-data quantity, quality, and coverage.Model confidence can be assessed through ensemble disagreement or distances between test samples and training neighbors in learned spaces.
  • UQ of ML models: UQ of machine-learning models mainly targets epistemic uncertainty, while aleatory uncertainty can often be learned directly from data.The review compares Gaussian processes, Bayesian neural networks, ensemble methods, and related approaches for predictive uncertainty estimation.
  • UQ of dynamic-system models: Dynamic-system UQ propagates uncertainty to outputs or estimates model uncertainty from observations, with the review focusing on the latter.State-space models represent states as numerical functions of inputs and uncertain parameters, with residual errors capturing model-form discrepancies.
  • UQ of dynamic-system models: Recovering missing physics in the state-transition equation could improve extrapolation more effectively than quantifying uncertainty only in the measurement equation.The review notes that most existing methods address measurement or state-transition equations separately, while relatively few apply the KOH framework to dynamic systems.

2.2 Optimization for digital twins (OPT)

Optimization in digital twins is divided into offline optimization before deployment and online optimization during operation. The review then considers techniques for applying optimization in deployed and pre-deployment settings.

  • Optimization categories: Offline optimization occurs before a digital twin is deployed, whereas online optimization occurs after deployment while the twin is operating.The paper uses this distinction to organize its discussion of optimization techniques for digital twins.

2.2.1 Optimization for sensor placement (offline)

Offline sensor placement optimization determines where sensors collect informative data for probabilistic model updating, while accounting for aleatory uncertainty and sensor-specific objectives. The review covers cost functions, computational strategies, and an example where optimal placement concentrates posterior damage estimates.

  • Sensor locations strongly affect data quality, inferred information, the P2V connection, and digital-twin performance.
  • Aleatory uncertainty is irreducible, so placement optimization should account for operational variability when designing sensor networks.
  • Cost functions: Information-gain objectives quantify uncertainty reduction using measures including Fisher’s information matrix, Kullback–Leibler divergence, and other f-divergences.
  • Cost functions: Probability-of-detection objectives minimize type I and type II errors, corresponding respectively to false alarms and missed detections.
  • The reviewed sensor-placement metrics are representative rather than exhaustive, and uncertainty-aware cost evaluation creates substantial computational challenges.
  • Example: 2021’s miter-gate example found that optimal sensor placement produced a much more concentrated posterior damage estimate than non-optimal placement.

2.2.2 Optimization for physical system modeling (offline)

Offline modeling optimization calibrates uncertain digital-model parameters before deployment, then supports feasible online updating by keeping the digital state and online-updated parameter subset small. The review contrasts Bayesian and optimization-based calibration and emphasizes validation.

  • A digital state uses state variables x and model parameters ˜θ=[λ, θ], with only the small subset θ updated online while λ remains fixed.
  • Offline calibration bridges an initial digital model and its physical counterpart using experimental data and Bayesian or optimization-based methods.
  • Bayesian calibration estimates parameter posteriors, while optimization-based calibration maximizes or minimizes a calibration metric.
  • Accounting for model discrepancy during Bayesian calibration can improve prediction accuracy and make the posterior estimate closer to the true parameter value.
  • Least squares is widely used and easy to implement but sensitive to outliers, whereas maximum likelihood may perform poorly with small calibration datasets.
  • Validation: Model validation is required after calibration to quantify agreement between digital-model predictions and experimental observations for intended uses.

2.2.3 Optimization for predictive decision making (online)

Online optimization uses digital-twin predictions and continual physical-system data to support real-time control, mission planning, and individualized maintenance decisions. Examples include melt-pool regulation, terrain-aware vehicle routing, and unit-specific retirement or maintenance scheduling.

  • Real-time requirements depend on the application’s system timescale, and acceptable delay may replace stricter computational-speed requirements for very high-rate systems.
  • Real-time process control: Layerwise powder-bed-fusion control trains a melt-pool prediction model from prior-build process commands and monitoring data, then optimizes scan parameters online.
  • Real-time process control: Laser power is selected as the sole optimization control because changing scan speed or scan path requires significant computing effort.
  • Mission planning: Off-road vehicle digital twins predict mobility across terrain and generate probabilistic maps identifying where autonomous vehicles can go or cannot go.
  • Predictive maintenance scheduling: Predictive maintenance requires a prognostic model for remaining useful life and an optimization model combining that estimate with maintenance preferences.
  • Predictive maintenance scheduling: Digital twins can use automatic prognostic-to-optimization data flow to tailor maintenance timing to an individual unit’s current and future health.

2.2.4 Summary of commonly used optimization methods in digital twins

Digital-twin optimization methods span modeling, P2V, and V2P applications, with method choice shaped by computational cost, data requirements, scalability, and real-time needs. Evolutionary, reinforcement-learning, Bayesian, and search-based methods each have distinct trade-offs.

  • Evolutionary methods can optimize physical-system modeling and V2P tasks such as process control, mission planning, and maintenance scheduling.
  • Evolutionary optimization usually requires a very high number of function evaluations, limiting its suitability when evaluations are expensive.
  • Reinforcement-learning optimization has emerged as a promising global approach but requires high training-data volumes and is complex to implement.
  • Bayesian optimization with Gaussian-process regression is widely used but may suffer from dimensionality-related scalability limitations.
  • Mission and path planning use search algorithms such as A* and RRT* for V2P applications.

3 Case study: a battery digital twin

The case study illustrates a battery digital twin implementation. Code and preprocessed data for reproducing its results and figures are available on GitHub.

  • The case study implements a battery digital twin.
  • The study provides code for generating all case-study results and figures.
  • Preprocessed case-study data are available on GitHub.

3.1 Background

Battery degradation depends on operating history, making accurate remaining-useful-life prediction important for maintenance, replacement, and retirement decisions. The review also distinguishes prognostic models from complete digital twins, which must include optimization and control.

  • Li-ion batteries lose capacity through irreversible internal electrochemical changes after many operating cycles.
  • Accurate remaining-useful-life predictions can inform decisions about cell maintenance, replacement, or retirement.
  • Cells with similar first-life applications can have different present health and remaining capacity because usage history shapes degradation.
  • Many reported digital-twin models cover only three or four of the five dimensions defined in the review.
  • A prognostic model alone is incomplete because its data flow ends after predicting physical-system RUL without optimization.
  • The proposed proof-of-concept battery twin combines particle-filter prognostics with multi-attribute optimization to determine when a Li-ion cell should leave first-life use.

3.2 Methods

The methods combine offline particle-filter calibration, online battery prognostics, and multi-attribute utility optimization. The battery model represents diverse capacity-fade behavior while the optimization accounts for user preferences and operating objectives.

  • Battery digital twin framework: Previously collected run-to-failure data are used offline to optimize the particle-filter parameters before online operation.
  • Battery degradation model: Battery capacity degradation is affected by charging and discharging rate, depth of discharge, time, and ambient temperature.
  • Particle filter prognostic model: The review describes recursive filtering as a way to estimate capacity-model parameters and generate probabilistic RUL distributions with uncertainty.
  • Particle filter prognostic model: Particle filtering is selected because it can switch between multiple models and estimate a non-parametric distribution.
  • Particle filter prognostic model: The case study uses a two-parameter power-law model for the diverse capacity-fade trends in its open-source dataset.
  • Multi-attribute utility optimization: Multi-attribute utility theory maps differently scaled objectives to a common range for evaluation and optimization.
  • Multi-attribute utility optimization: The retirement decision maximizes utility over total discharge ampere-hours and mean time between charges, with cycle number as the decision variable.

3.3 Results and discussion

The battery digital twin combines particle-filter RUL prediction with utility-based optimization to determine when Li-ion cells should retire from first-life use. Tests show systematic RUL underestimation under distribution shift, while optimized retirement depends mainly on capacity-loss rate and utility trade-offs.

  • Open-source battery dataset: The case study uses 124 LFP/graphite cells with varied two-step fast-charging protocols, split into training, primary-test, and secondary-test datasets.The training set tunes prognostic hyperparameters before online testing on the two test datasets.
  • RUL prediction results: The particle filter projects future capacity trajectories from many particles and derives an empirical end-of-life distribution.Capacity measurements and median projected capacity are compared to characterize future cell behavior.
  • RUL prediction results: The particle filter mostly underestimates RUL, with stronger underestimation for cells exceeding 2000 cycles.Linear extrapolation of late-life capacity loss and median initialization of state parameters contribute to this pattern.
  • RUL prediction results: The test datasets have a median lifetime of 1125 cycles versus 750 cycles for training, creating distribution shift that drives general RUL underestimation.Untuned measurement and process noise may also limit state variability and slow convergence toward true RUL.
  • Defining the utility functions: The optimization combines utility functions for Ah throughput and mean time between charges to encode first-life retirement preferences.Exponential utility functions map attributes to comparable perceived values, while scaling parameters set utility limits.
  • Analyzing the optimization results: Optimal retirement is most closely related to observed capacity-loss rate and typically occurs shortly after the capacity-fade knee point.Longer-lived cells can be retired earlier when the Ah-throughput utility is nearly constant across possible retirement cycles.
  • Analyzing the optimization results: For longer-lived cells, underestimated capacity trajectories move optimal replacement closer to the current cycle, whereas projected trajectory shape otherwise has limited effect.The reported retirement decisions remain consistent when projected and measured curve shapes are fairly similar.

3.4 Case study conclusion and ideas for future research

The battery digital twin integrated particle-filter prognostics with multiattribute utility optimization to estimate cell retirement timing. The case study also identifies model, utility-function, preference, and probabilistic-decision limitations for future work.

  • The case study created a battery digital twin that optimizes when a Li-ion cell should retire from first-life use.
  • The particle-filter prognostic model accurately predicted remaining useful life for cells with varying lifetimes.
  • Future work should examine multi-model particle filters, broader utility functions, second-life preferences, and probabilistic replacement intervals rather than deterministic point estimates.The case study omitted second-life attributes and used a single power-law capacity-fade model; probabilistic utility evaluation could provide confidence levels and intervals.
  • The multiattribute utility model provided advance notice of optimal retirement using projected future capacity from the particle filter.
  • The optimization model struggled to produce meaningful replacement predictions for cells with much longer lifetimes.

4 Demonstration and open source

The review demonstrates digital-twin applications across industries and catalogs open-source tools and datasets supporting modeling, process mining, control, data streaming, and implementation. It argues that broader sharing of tools, data, and practices can accelerate industry-scale adoption.

  • 4.1 Industry-scale demonstration of digital twin: $44 million and $10 million savings were associated with GE Aviation’s reduced anomaly-detection time for turbine-engine life-cycle management and flight-pattern and maintenance optimization.
  • 4.1 Industry-scale demonstration of digital twin: Digital twins have been applied across aerospace, commercial, defense, automotive, healthcare, and industrial settings to support design, quality, maintenance, and optimization.
  • 4.1 Industry-scale demonstration of digital twin: 40% improvement in first-time quality was reported by Boeing for parts and systems managed with operational digital twins.
  • Open-source tools and data: The review identifies publicly available tools and datasets spanning modeling, process mining, control, data streaming, and digital-twin projects.Examples include Chrono for multiphysics simulation and the Bosch Production Line Performance Dataset for quality-control prediction.
  • Open-source tools and data: Open-source tools and datasets could support production-level digital-twin implementations in academia and industry.
  • Open-source tools and data: Industry-scale adoption is expected to accelerate through sharing tools, data, best practices, tutorials, implementation guides, and training materials across sectors.

5 Perspectives on UQ and optimization for digital twins

The paper’s perspectives span lifecycle digital twins, sustainability applications, battery repurposing, and uncertainty-aware decision making. It highlights both opportunities for optimization and constraints from limited data and delayed validation.

  • Digital twins for structural life cycle management: Digital representations can support structural visualization, state awareness, future-performance prediction, and life-cycle management decisions such as maintenance and repair.These representations may combine surveys, IoT streams, learning methods, physical degradation knowledge, uncertainty quantification, and utility or cost models.
  • Digital twins for sustainability: Digital twin research covers development, manufacturing, and service more often than disposal, leaving reuse, remanufacturing, and recycling comparatively underexplored.The review calls for more research connecting digital twins with wide-scale remanufacturing adoption.
  • Digital twins for sustainability: Remanufacturing faces limited historical and representative run-to-failure data for returned products, constraining fault diagnosis and remaining-useful-life prediction.The paper identifies physics-informed machine learning as one approach for addressing these data challenges.
  • Battery repurposing: Battery repurposing requires rapid state-of-health and remaining-useful-life estimates to assess the economic and technical viability of second-life applications.A Battery Passport is presented as a digital twin containing relevant battery information from resource extraction and manufacturing through repurposing or recycling.
  • UQ of digital twins: Validation is a lagging indicator because a specimen-specific digital twin can be validated only with data from the specimen’s future use.This differs from validation of a general-purpose prediction model, which can use experiments on a population of specimens or realizations.
  • UQ of digital twins: Current models incur additional uncertainty before future validation, and extrapolation becomes more difficult when aging systems operate beyond design life or outside design conditions.The paper notes that uncertainty is expected to decrease over time only when operating conditions remain similar.

6 Conclusion

The conclusion positions uncertainty quantification and optimization as central to improving digital-twin usefulness, then illustrates their combination in a battery application. The battery twin determines when a cell should leave its first-life application and provides lifecycle status information for practitioners.

  • The review covers uncertainty in physical measurements and virtual outputs, along with optimization techniques intended to improve digital-twin accuracy, reduce uncertainty, and increase usefulness.
  • The battery digital twin combines degradation modeling and optimization to determine the optimal retirement time of a battery cell from its first-life application.
  • The battery case provides practitioners with actionable information about a cell’s status across its overall life cycle.
  • The paper presents digital-twin modeling as a developing basis for managing, maintaining, and retiring high-value assets.

Authors’ contributions

The authors divided responsibility across the review’s conceptual, literature, geometric, physics-based, data-driven, machine-learning, and system-modeling components.

  • The authors assigned distinct responsibilities for the review concept, literature review, geometric and physics-based modeling, data-driven modeling, physics-informed machine learning, and system modeling.
Loading 2208.12904v1…