Source-linked AI summary

A Hybrid Science-Guided Machine Learning Approach for Modeling and Optimizing Chemical Processes

Niket Sharma, Y. A. Liu

arXiv:2112.01475v2cs.LG

TL;DR

The paper addresses how to combine scientific knowledge and data analytics when data-based models can be scientifically inconsistent and science-based models have knowledge or parameter-estimation gaps. It reviews and classifies hybrid SGML methods in which ML complements science or scientific principles complement ML, covering applications and polymer-process examples. The review recommends hybrid models over standalone ML for process development because they better support extrapolation beyond tested conditions.

  • Problem

    Data-based models can be scientifically inconsistent and data-intensive, while science-based models may have knowledge gaps and multiple parameters that are difficult to estimate.

  • Method

    The paper reviews and systematically classifies hybrid SGML methods in which ML complements science-based models or scientific knowledge complements ML models.

  • Results

    The review recommends hybrid models over standalone ML for process development because they are better at extrapolating beyond experimentally tested process conditions.

  • Takeaways & Limitations

    Hybrid SGML provides a framework for combining scientific principles and data analytics across bioprocessing and chemical engineering applications.

  • Takeaways & Limitations

    A data-based model may have similar accuracy to a hybrid model but can produce scientifically inconsistent predictions beyond the operating data used for training.

Abstract

from arXiv · show

This study presents a broad perspective of hybrid process modeling and optimization combining the scientific knowledge and data analytics in bioprocessing and chemical engineering with a science-guided machine learning (SGML) approach. We divide the approach into two major categories. The first refers to the case where a data-based ML model compliments and makes the first-principle science-based model more accurate in prediction, and the second corresponds to the case where scientific knowledge helps make the ML model more scientifically consistent. We present a detailed review of scientific and engineering literature relating to the hybrid SGML approach, and propose a systematic classification of hybrid SGML models. For applying ML to improve science-based models, we present expositions of the sub-categories of direct serial and parallel hybrid modeling and their combinations, inverse modeling, reduced-order modeling, quantifying uncertainty in the process and even discovering governing equations of the process model. For applying scientific principles to improve ML models, we discuss the sub-categories of science-guided design, learning and refinement. For each sub-category, we identify its requirements, advantages and limitations, together with their published and potential areas of applications in bioprocessing and chemical engineering.We also present several examples to illustrate different hybrid SGML methodologies for modeling polymer processes.

1 | INTRODUCTION

The paper motivates hybrid SGML because data-based models can overfit, require substantial data, and produce scientifically inconsistent predictions, while scientific models may leave knowledge gaps. It reviews and classifies approaches in which ML complements science-based models or scientific knowledge complements ML models.

  • Data-based ML models can overfit, require more data, and produce scientifically inconsistent results outside observed operating conditions.
  • Hybrid SGML integrates science-based knowledge with data-based knowledge to support accurate and scientifically consistent prediction.
  • The review classifies hybrid approaches into ML complementing science and science complementing ML, with subcategories organized by methodology and objective.
  • The paper surveys requirements, strengths, limitations, applications, and illustrative chemical-manufacturing examples across both categories.
  • Its contributions include broader SGML coverage, methodology- and objective-based classification, underexplored research themes, and polymer-process illustrations including inverse modeling and science-guided loss.

ENGINEERING

Prior work applies hybrid modeling across bioprocessing and chemical engineering for prediction, control, optimization, design, scale-up, and model reduction. The literature includes direct hybrid structures, gray-box and semi-parametric models, reduced-order models, and broader AI–domain-knowledge integrations, while emphasizing that knowledge integration does not automatically improve results.

  • Integrated models can improve interpolation and extrapolation accuracy, interpretability, and training-data efficiency relative to standalone data-based models.
  • Examples include sparse-data fermentation prediction, cumene-process optimization, fed-batch bioreactors, industrial distillation, and applications in digital twins and asset optimization.
  • Common structures include series and parallel direct hybrids, gray-box models, semi-parametric models, surrogate or reduced-order models, and alternate structures.
  • Reported applications span monitoring, control, optimization, process development, scale-up, process design, experiment design, predictive maintenance, and model reduction.
  • Hybrid modeling does not automatically produce better results; assumptions, parameter-estimation time, accuracy, and potentially biased scientific knowledge constrain strategy selection.
  • Hybrid modeling has been applied to bioprocessing, chemical, oil and gas, polymer, and separation processes for scientifically consistent predictions.

MODELS

The paper broadens hybrid modeling beyond direct science–data combinations by organizing SGML applications according to whether ML complements science or science complements ML. It uses methodology- and objective-based categories to identify underexplored opportunities for process improvement and optimization.

  • Most existing applications focus on directly combining science-based and data-based models, whereas this paper presents a broader SGML perspective.
  • The classification has two major categories: ML complements science, and science complements ML.
  • Subcategories are organized according to hybrid-modeling methodologies and objectives and are illustrated in Figure 1.
  • The paper identifies relatively unexplored SGML examples with potential for process improvement and optimization.

4 | ML COMPLEMENTS SCIENCE

The section classifies direct hybrid models that combine science-based and ML components in parallel, series, or series-parallel configurations to improve prediction accuracy and extrapolation. Applications span polymerization, bioprocess control, process optimization, predictive maintenance, CFD-based simulation, and inverse modeling.

  • Direct hybrid modeling: Direct hybrid models combine science-based and ML outputs in series, parallel, or series-parallel configurations to improve dependent-variable prediction accuracy.The section identifies direct hybrid modeling as the most widely used hybrid approach in bioprocessing and chemical engineering.
  • Parallel direct hybrid model: Residual hybrids train ML models to estimate time-dependent errors between plant data and science-based outputs, then add those residuals to the model predictions.This correction improves accuracy over the non-residual parallel configuration.
  • Parallel direct hybrid model: In batch polymerization, recurrent neural networks corrected simplified kinetic and balance models for monomer conversion, MWN, and MWW, supporting batch control and optimization.The approach addresses omitted effects such as the gel effect at high monomer conversion and uses time-dependent data for long-range prediction.
  • Series direct hybrid model: Series hybrids use science-based calculations to augment ML inputs or estimate science-model parameters, improving production prediction and extrapolation.Examples include crude-distillation quality analysis, industrial reactor prediction, and CFD-guided process simulation and optimization.
  • Application to polymer manufacturing: For polymer manufacturing, combined hybrid predictions achieved RMSE 0.21 and matched plant data better than a first-principles dynamic simulation alone.The comparison concerns melt-index prediction during plant operation with grade transitions.
  • Inverse modeling: Inverse hybrid modeling uses process-variable and quality-target data to predict conditions for product formation, achieving a success rate of nearly 90%.Inverting the ML model also revealed hypotheses about conditions associated with successful product formation and reduced experimental burden in pharmaceutical development.

4.3 | An Application of Inverse Modeling to Polymer Manufacturing

Inverse modeling uses science-based process simulations and machine learning to predict operating conditions from desired polymer quality targets. Reduced-order models similarly support efficient process optimization, soft sensing, and variable screening.

  • Inverse modeling: The inverse workflow simulates polymer quality across operating conditions, then trains an ensemble regressor mapping quality targets to input-stream flow rates.Given desired quality targets, the trained model predicts operating conditions for a new polymer grade.
  • Inverse modeling: RMSE = 0.9 for hydrogen feed-flow prediction against actual plant data demonstrates accurate inverse modeling.The model predicts operating conditions for producing a new polymer grade from targets such as melt index, density, polydispersity, and production rate.
  • Reduced-order modeling: Reduced-order models use dimensional reduction, residual learning, or ML surrogates to approximate complex process models more cheaply.These models can support online deployment, feasibility analysis, dynamic optimization, and feedback-control design.
  • Reduced-order modeling: A compartmentalized dynamic model combined with neural networks reduced the differential-equation system size by 90% while keeping product-purity error below 1 ppm.The comparison is against a full-order stage-by-stage model.
  • Reduced-order modeling: Science-based simulations generate varied training data for ML soft sensors, improving coverage beyond what is available from steady plant operation.In the HYPOL polypropylene example, the model predicts melt index and identifies hydrogen flow rate and fourth-reactor temperature as important variables for optimization.

Manufacturing

The paper illustrates uncertainty quantification and scientific-law discovery as hybrid SGML applications. Prediction intervals expose when melt-index predictions are uncertain, while ML can help identify or refine governing physical laws.

  • Uncertainty quantification: RMSE values of 1.2–1.5 versus melt-index standard deviation 5.1 quantify prediction uncertainty for the slurry HDPE process.The model uses prediction intervals to characterize the range of possible melt-index predictions.
  • Uncertainty quantification: A larger prediction interval before 100 hours indicates higher uncertainty during the period of greater melt-index change.The interval is defined by the 5th and 95th quantiles, representing a 90% prediction interval.
  • Uncertainty quantification: Uncertainty quantification supports process decisions by exposing the model’s error estimate rather than providing predictions alone.The paper describes quantile-regression gradient boosting as the method for estimating prediction intervals.
  • Discovering scientific laws: ML can discover functional forms or parameters of physical and chemistry laws for incorporation into science-based process models.Examples include phase-equilibrium relations, governing PDEs, and chemical laws recovered from data.
  • Discovering scientific laws: Discovered laws can improve first-principles model accuracy while reducing model complexity.The paper frames data-driven discovery as a route to improving scientific process models.

5 | SCIENCE COMPLIMENTS ML

Science-guided ML incorporates scientific knowledge into model architecture, learning objectives, initialization, and post-processing. The reviewed examples aim to improve consistency, interpretability, extrapolation, or efficiency while retaining predictive capability.

  • Science-guided design and learning: Scientific knowledge can guide ML architecture, learning, initialization, and post-processing to improve generalization and reduce scientific inconsistency.The paper organizes these mechanisms as science-guided design, learning, and refinement.
  • Science-guided design: Fuzzy ANNs connect process knowledge to network rules and weights, supporting scientific consistency, lower computational complexity, interpretability, and extrapolation.The paper highlights applications in process control and bioprocess monitoring.
  • Science-guided learning: A science-guided loss combines prediction error with a science-based loss weighted by λ.The prediction term compares Y_true and Y_pred, while the science-based term penalizes scientifically inconsistent predictions.
  • Science-guided refinement: Science-guided initialization can improve training by providing initial parameters from science-based models and helping avoid local minima.The paper relates this approach to transfer learning and process-model migration.

PROCESSES

Hybrid SGML models offer opportunities for extrapolation, process development, optimization, monitoring, and control, but their performance depends on scientific-model quality, data, computation, and domain expertise. The paper summarizes these trade-offs across model classes.

  • Limitations: First-principles assumptions and inaccurate scientific models can propagate error into hybrid-model predictions.The paper emphasizes that scientific-model accuracy is important for reliable hybrid modeling.
  • Limitations: Applying hybrid SGML can require expertise spanning both domain science and machine learning, alongside difficult data preparation and feature engineering.These requirements can increase complexity relative to standalone neural networks.
  • Applications: Hybrid SGML models are useful for extrapolation beyond the operating range and may support process development.The paper also identifies fault diagnosis and anomaly detection as potential application areas.
  • Limitations: Inverse modeling may have non-unique solutions, increased computational demands, and lower generality.These limitations constrain the interpretation and deployment of inverse approaches.
  • Model-class trade-offs: Reduced-order models enable fast online deployment and lower complexity but can introduce bias and remain limited by science-based-model accuracy.The trade-off is summarized for reduced-order SGML models in the paper’s classification table.

7 | CONCLUSION

The paper presents a broad hybrid SGML review and classification, distinguishing ML-enhanced science-based models from science-guided ML models. It illustrates applications in industrial polymer and chemical process improvement and concludes that hybrid models generally outperform standalone ML for extrapolation.

  • It reviews hybrid SGML methodologies and identifies underexplored opportunities for chemical process modeling.
  • The paper classifies hybrid SGML into ML-enhanced first-principles models and scientifically consistent ML models.
  • The approach is illustrated through applications in industrial polymer and chemical process improvement.
  • Hybrid models generally outperform standalone ML for extrapolation, whereas standalone ML can be adequate under steady operating conditions.
Loading 2112.01475v2…