Source-linked AI summary

Turbulence Modeling in the Age of Data

Karthik Duraisamy, Gianluca Iaccarino, Heng Xiao

arXiv:1804.00183v3physics.flu-dynphysics.comp-ph

TL;DR

Turbulence models face challenges in assessing predictive capability and quantifying model inadequacies as data-driven methods expand. This review synthesizes physical constraints, statistical inference, and machine learning, concluding that approaches grounded in turbulence knowledge can improve predictive models.

  • Problem

    Assessing turbulence-closure predictiveness and quantifying model inadequacies require more consistent and transparent uncertainty analysis.

  • Method

    The review surveys physical-constraint uncertainty bounds, statistical inference for coefficients and discrepancy, and machine learning for turbulence-model improvement.

  • Results

    Data-driven corrections produced improved Reynolds-stress anisotropy and mean-velocity predictions in canonical flows, while limited measurements also improved airfoil-flow predictions.

  • Takeaways & Limitations

    Useful predictive turbulence models can result when data-driven approaches exploit foundational turbulence knowledge and physical constraints.

  • Takeaways & Limitations

    Data-driven models must address limited problem relevance of available data and consistency between learning environments such as DNS and deployment environments such as RANS.

Abstract

from arXiv · show

Data from experiments and direct simulations of turbulence have historically been used to calibrate simple engineering models such as those based on the Reynolds-averaged Navier--Stokes (RANS) equations. In the past few years, with the availability of large and diverse datasets, researchers have begun to explore methods to systematically inform turbulence models with data, with the goal of quantifying and reducing model uncertainties. This review surveys recent developments in bounding uncertainties in RANS models via physical constraints, in adopting statistical inference to characterize model coefficients and estimate discrepancy, and in using machine learning to improve turbulence models. Key principles, achievements and challenges are discussed. A central perspective advocated in this review is that by exploiting foundational knowledge in turbulence modeling and physical constraints, data-driven approaches can yield useful predictive models.

1. Introduction

Turbulence is important across engineering applications, but its broad range of scales and chaotic nature makes representation challenging. Recent advances in computing and data science have opened new perspectives for turbulence modeling, including uncertainty-aware approaches applicable to RANS and, in some respects, LES.

  • Engineering motivation: Turbulence knowledge supports wind-turbine performance, fuel/air mixing and emissions reduction in engines, and reduced aircraft fuel consumption.These applications depend on incoming-flow and boundary-layer turbulence, vigorous mixing, or delayed boundary-layer transition.
  • Modeling challenge: Representing turbulent motions is challenging because they span broad spatial and temporal scales and exhibit strong chaotic behavior.The passage situates this challenge within decades of theoretical and computational turbulence research.
  • Data-driven perspective: Recent advances in data science are offering new perspectives to the classical field of turbulence modeling.The article title reflects this perspective and its connection to contemporary data-science developments.
  • Modeling approaches: LES represents part of the active scales while modeling unresolved motions, and ideas developed for RANS may also apply to LES.LES is gaining popularity in industrial applications with relatively small Reynolds numbers, but it still involves modeling assumptions.

Lexicon of data-driven modeling

The lexicon defines a data-driven model through its computational model, data and uncertainty, model output, and discrepancy. The framework also represents discrepancy using features informed by prior knowledge, constraints, or data, with interest in predicting quantities of interest.

  • Model elements: M denotes the computational model, a function of independent variables w and operators P, with parameters c targeted by data modeling.The operators may be algebraic or differential.
  • Model elements: θ denotes the data with uncertainty ϵθ, while o is the model output corresponding to θ.The data may be accompanied by a quality estimate.
  • Model elements: δ denotes discrepancy between the model and data; it is unknown, generally depends on the model, and is described using θ and features η.Features η can derive from prior knowledge, constraints, or directly from data.
  • Predictive objective: The general data-driven framework is used to predict quantities of interest q(f_M).The supplied passage introduces this predictive objective without providing the full model equation.

2. Turbulence closures and uncertainties

RANS closures introduce four layers of simplification, from irrecoverable ensemble-averaging uncertainty to assumptions about Reynolds-stress representation, functional forms, and coefficients. These layers complicate predictive assessment and require transparent quantification of model inadequacies.

  • Implications: Together, the four modeling layers make predictive credibility difficult to assess and require careful, transparent procedures for quantifying and reducing model inadequacies.They also create inconsistency when different modeling strategies are compared.
  • L1: Ensemble-averaging uncertainty: Ensemble-averaging loses information about microscopic velocity realizations, creating irrecoverable uncertainty regardless of turbulence-model sophistication.Different realizations compatible with one averaged field can evolve differently, producing hysteresislike behavior.
  • L2: Reynolds-stress representation: Reynolds-stress closures represent the macroscopic state using selected averaged variables, introducing uncertainty in the functional and operational representation.Examples include one-point or two-point closures, linear eddy viscosity models, and algebraic stress models.
  • L3: Functional-form uncertainty: After selecting independent variables, researchers postulate algebraic or differential functional forms to represent physical processes and modeling assumptions.One- and two-equation models are common, with additional terms representing sensitivities such as near-wall dynamics and rotational corrections.
  • L4: Coefficient uncertainty: Given a complete model structure and functional form, coefficients calibrate the relative importance of closure contributions, creating a further uncertainty layer.Coefficients may follow asymptotic consistency or empirical evidence; choosing Cµ in two-equation linear eddy viscosity models is a classical example.

Uncertainty Quantification

Uncertainty quantification addresses uncertainty in numerical predictions by distinguishing variability in simulation inputs from limitations in physics models. It identifies and probabilistically describes uncertainty sources, then propagates them through the model to predict quantities of interest.

  • Uncertainty sources: UQ distinguishes aleatory uncertainty from epistemic and model-form uncertainty in numerical predictions.Aleatory uncertainty reflects imprecision or natural variability in real-world simulation inputs, whereas epistemic and model-form uncertainties arise from intrinsic physics-model limitations.
  • UQ workflow: UQ first identifies uncertainty sources and introduces an appropriate probabilistic description, ϵ.The described uncertainty is then propagated through the model M to obtain predictions of a quantity of interest q(M, ϵ).
  • UQ workflow: Propagation through the model is typically computationally intensive, motivating the development of efficient UQ strategies.These strategies have received considerable attention in the last decade.

3. Models, data and calibration

This section describes how experimental and DNS data inform turbulence-model construction, from simple coefficient calibration to probabilistic inference and data-focused machine-learning approaches. It emphasizes that increasingly large datasets enable models to account for measurement uncertainty and model–data discrepancy.

  • Data sources and calibration: Experimental turbulence measurements have constrained model constants, including the (L4) constants c using measured isotropic-turbulence decay rates.The Comte-Bellot and Corrsin (1966) experiments provided the decay-rate observations used for this constraint.
  • Data sources and calibration: DNS databases have become an additional source of modeling insight, supported by sustained community efforts to gather and archive turbulence datasets.The Center for Turbulence Research Summer Program was established in 1987 to study turbulence using numerical simulation databases.
  • Simple calibration: Simple calibration selects similar experimental configurations, typically ignores measurement uncertainty, and treats model coefficients c as the dominant uncertainty source.Predictive accuracy is judged by the difference in q obtained with c or ˜cq, contributing to many model variants and difficulty assessing predictive capabilities.
  • Statistical inference: Statistical inference extends calibration by incorporating experimental uncertainty, multi-source data, and model–data discrepancy in a Bayesian probabilistic formulation.The resulting calibrated stochastic model includes parameter uncertainty, discrepancy δ, and measurement error ϵθ; discrepancy priors may use Gaussian random fields with estimated parameters.
  • Data-focused approaches: Efficient inference algorithms enable assimilation of large DNS datasets and shift emphasis from traditional model M toward discrepancy δ, with machine-learning representations increasingly used.Different functional representations of δ are available as data-focused approaches develop.

Statistical inversion

Statistical inversion calibrates model parameters by reconciling uncertain data with model outputs, using Bayesian inference or least squares. Calibrated models can propagate stochastic information into predictions while embedding calibration within the model rather than treating the input–prediction mapping as physics-agnostic.

  • Statistical inversion: Statistical inversion identifies model parameters c by minimizing the difference between uncertain data θ and the corresponding model output o(M(c)).In the Bayesian formulation, the posterior probability of c given θ combines prior information with the likelihood of consistency between the model and data.
  • Statistical inversion: Under Gaussian probability distributions, the maximum a posteriori estimate cMAP is obtained from a deterministic optimization problem involving observation and prior covariance matrices.Qθ and Qc denote the observation and prior covariance matrices, respectively.
  • Statistical inversion: Least Squares minimizes the discrepancy between θ and o(M(c)), while a γ-scaled regularization term improves inversion well-posedness and conditioning.With Gaussian distributions and γ ≡ QθQ−1c, the least-squares and maximum-a-posteriori solutions coincide: cLS = cMAP.
  • Statistical inversion: The review embeds calibration inside the model so propagated stochastic information gives predictions with uncertainty reflecting the calibrated model’s ability to represent the data.This contrasts with treating the entire input-to-prediction mapping as a physics-agnostic black box.

4. Quantifying uncertainties in RANS models

This section reviews interval and probabilistic strategies for quantifying uncertainty in RANS Reynolds-stress predictions, arising from closure assumptions and calibration. It also considers physical constraints, spatially varying discrepancies, parameter inference, and markers for potentially uncertain regions.

  • Sources of uncertainty: RANS prediction uncertainty arises from closure assumptions and calibration, motivating models that represent discrepancy through theoretical arguments or comparisons with existing data.The reviewed formulation is f_M = M + ϵ_M, with either theoretically derived discrepancy or ϵ_M = ϵ(θ) inferred from data comparisons.
  • Uncertainty descriptions: The section surveys interval and probabilistic descriptions of uncertainty in Reynolds stresses.Interval approaches seek bounds containing true answers, whereas probabilistic approaches characterize uncertainty through distributions.
  • Physics-constrained and probabilistic methods: Realizability-constrained eigenspace perturbations and maximum-entropy random matrices provide two Reynolds-stress uncertainty approaches that preserve physical admissibility.The eigenspace method perturbs stresses toward one-component (1C), two-component (2C), and three-component (3C) states; the random-matrix method models stresses on symmetric positive semi-definite 3 × 3 matrices with E[T] = τ_RANS.
  • Spatial variation: Uncertainty quantification must account for spatially varying Reynolds-stress discrepancy because its divergence enters the RANS equations.Existing approaches specify spatial fields for eigenvalue perturbations based on assumed limitations of RANS models.
  • Uncertainty localization: Markers identify regions where RANS predictions may be unreliable, including areas associated with inaccurate Reynolds-stress divergence or emerging negative eddy viscosity.Gorlé et al.’s analytical marker correlates with inaccurate Reynolds-stress-divergence predictions, while Ling and Templeton used DNS and RANS databases to identify likely linear eddy-viscosity-model failures.

5. Predictive modeling using data-driven techniques

This section reviews data-driven approaches that improve RANS prediction accuracy by deriving models with explicit discrepancy functions. It covers Bayesian calibration, discrepancy-field inference, and the limited generalizability that motivates combining turbulence modeling with learning strategies.

  • Section focus: Data-driven prediction approaches seek to derive f_M = M(θ) while explicitly introducing a discrepancy function δ.The section shifts from quantifying model uncertainty toward improving overall prediction accuracy using data.
  • Statistical inference: Bayesian inference used DNS data to assign posterior probability distributions to parameters of several turbulence models.Oliver and Moser (2011) and Cheung et al. (2011) were among the first to apply this approach to plane channel flows.
  • Statistical inference: Scenario averaging and related studies calibrated RANS coefficients, assessed competing closures, and incorporated diverse datasets from wall-bounded flows and jet-in-cross-flow.These approaches used statistical inference to construct posterior distributions based on data.
  • Discrepancy modeling: Discrepancy-based methods extended calibration by representing uncertainty as a Reynolds stress discrepancy tensor or stochastic turbulent-viscosity field used for prediction.Gaussian process models and limited measurements were used to infer discrepancy fields in channel-flow applications.
  • Limitations and machine learning: Spatially varying discrepancy fields inferred from geometry-specific data are not easily generalizable, motivating approaches that combine turbulence modeling, inference, uncertainty quantification, and learning.Scenario averaging offers some relief by incorporating evidence from different datasets, but the limitation led to further machine-learning developments.

Machine Learning

Machine learning offers flexible ways to map large datasets to quantities of interest and improve turbulence models, either directly or through corrections to existing models. Recent approaches emphasize physics-informed representations, invariant Reynolds-stress modeling, and combinations of statistical inference with machine learning.

  • Machine-learning foundations: Supervised learning constructs mappings to specified targets, whereas unsupervised learning discovers patterns and reduces data complexity.Clustering and dimension reduction are examples of unsupervised learning.
  • Neural networks: Neural networks approximate complex functions through compositions of nonlinear functions, while deep networks use intermediate layers to represent arbitrary complex functional forms.Training typically requires determining a very large number of calibration coefficients, supported by efficient algorithms on modern computer architectures.
  • Applications: Machine learning can operate as a black-box mapping or provide a posteriori corrections to existing turbulence models.Applications include supervised learning of perturbations to Barycentric coordinates and broader prediction of Reynolds-stress discrepancies.
  • Physics-informed modeling: Invariant Reynolds-stress models use tensor invariants and orientation representations such as Euler angles or unit quaternions to promote objectivity and rotational invariance.These representations were used in several machine-learning approaches to Reynolds-stress modeling.
  • Hybrid inference and learning: Combining statistical inference with machine learning extracts spatial model discrepancy from representative datasets, reconstructs its functional form, and embeds it as an RANS correction.For turbulent flow over airfoils, the approach produced convincing prediction improvements and could use limited experimental measurements such as lift coefficient data.

6. Challenges and perspectives

The section identifies unresolved challenges in selecting relevant data, quantifying information content, accounting for uncertainty, and avoiding spurious laws from finite datasets. It advocates holistic modeling that combines data with prior structures and physical constraints, while expecting data-driven models and explicit uncertainty estimates to expand.

  • What data to use?: Available experimental and DNS databases may have limited relevance to specific problems, making data information content critical to quantify.Calibration should indicate whether additional data are needed or overfitting may occur.
  • What data to use?: Formal design-of-experiments can guide further data collection, while data uncertainty should be incorporated during inference and learning and propagated to predictions.This supports reasonable expectations about prediction accuracy.
  • Challenges and perspectives: With large but finite datasets, machine learning alone may discover problem-specific or spurious laws with limited predictive value.Model choices therefore depend on confidence in data, prior structures, physical constraints, and the model’s purpose.
  • Challenges and perspectives: Data-driven models are expected to grow as algorithms and computer architectures advance, influencing parameters, operators, and discrepancy.The recommended resulting predictions should include explicit uncertainty estimates.

DISCLOSURE STATEMENT

The authors report no known affiliations, memberships, funding, or financial holdings that might be perceived as affecting the review’s objectivity.

  • The authors disclose no affiliations, memberships, funding, or financial holdings perceived as affecting the review’s objectivity.
Loading 1804.00183v3…