Source-linked AI summary
Theory-guided Data Science: A New Paradigm for Scientific Discovery from Data
Anuj Karpatne, Gowtham Atluri, James Faghmous, Michael Steinbach, Arindam Banerjee, Auroop Ganguly, Shashi Shekhar, Nagiza Samatova, Vipin Kumar
TL;DR
Scientific data science models have limited applicability to complex physical phenomena because data can be scarce or misleading and black-box predictions may lack mechanistic interpretability. The paper formally conceptualizes theory-guided data science, presents a taxonomy and illustrative integration approaches, and concludes that TGDS offers a foundation for combining scientific knowledge with data-driven models while requiring further formalization.
Problem
Complex scientific problems combine limited representative data, intricate physical relationships, and a need for interpretable mechanisms that black-box data science alone does not provide.
Method
The paper formalizes TGDS, develops a taxonomy of ways to integrate scientific knowledge with data science, and illustrates the approaches across scientific disciplines.
Results
TGDS approaches span strict physical-consistency enforcement through model design or constraints and relaxed knowledge use through priors or regularization, with examples across diverse applications.
Takeaways & Limitations
TGDS provides a foundation for developing generalizable and scientifically interpretable models by combining theory-based knowledge with data-driven learning.
Takeaways & Limitations
Further progress requires innovations for representing and ingesting scientific knowledge across disciplines with differing granularity, completeness, and uncertainty.
Abstract
from arXiv · showhide
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of data science models in enabling scientific discovery. The overarching vision of TGDS is to introduce scientific consistency as an essential component for learning generalizable models. Further, by producing scientifically interpretable models, TGDS aims to advance our scientific understanding by discovering novel domain insights. Indeed, the paradigm of TGDS has started to gain prominence in a number of scientific disciplines such as turbulence modeling, material discovery, quantum chemistry, bio-medical science, bio-marker discovery, climate science, and hydrology. In this paper, we formally conceptualize the paradigm of TGDS and present a taxonomy of research themes in TGDS. We describe several approaches for integrating domain knowledge in different research themes using illustrative examples from different disciplines. We also highlight some of the promising avenues of novel research for realizing the full potential of theory-guided data science.
1 INTRODUCTION
Scientific data science offers new opportunities for discovery, but black-box methods often struggle with limited, complex data and cannot by themselves provide mechanistic understanding. TGDS is introduced to integrate scientific knowledge with data-driven models across disciplines.
- Motivation: Scientific data collection and computational analysis are expanding the role of data science from simple analysis toward knowledge discovery.The paper notes applications ranging from particle detection to frameworks in bioinformatics and climate science.
- Motivation: Black-box data science has had limited success in scientific domains, as illustrated by failures of theory-agnostic models such as Google Flu Trends.The example demonstrates the risks of relying on correlations without adequate scientific grounding.
- Scientific challenges: Scientific problems are often under-constrained because representative training samples are scarce while physical variables exhibit complex, non-stationary behavior.These conditions can make standard generalizability assessments misleading and allow spurious relationships to appear reliable.
- Scientific challenges: Scientific discovery requires interpretable theories and hypotheses that explain physical cause-effect mechanisms, not merely accurate predictive models.Interpretability can also help guard against spurious, non-generalizable patterns, especially in high-risk settings such as healthcare.
- TGDS: Theory-guided data science integrates scientific knowledge with data science to support discovery across climate science, turbulence modeling, materials, quantum chemistry, and other fields.The paper presents TGDS as a formally conceptualized paradigm and illustrates approaches using diverse applications.
2 THEORY-GUIDED DATA SCIENCE
TGDS addresses the limitations of data-only and theory-only approaches by combining scientific knowledge with data-driven learning. Its central objective is to favor models that balance accuracy, simplicity, and physical consistency while improving generalizability and interpretability.
- Motivation: Theory-based models encode known cause-effect relationships, whereas data science models use observations; each relies primarily on one source of scientific information.The paper places them at opposite ends of a knowledge-discovery continuum.
- Motivation: Theory-only models can perform poorly when complex processes are incompletely understood and require simplifying assumptions.The paper illustrates this problem with hydrological modeling of difficult-to-measure, nonlinear subsurface flow.
- Motivation: Data-only models may fail when available data cannot represent complex hypothesis spaces and when associative patterns do not reveal causative relationships.These limitations make neither data-only nor theory-only modeling sufficient for complex scientific knowledge discovery.
- TGDS framework: TGDS blends theory and data to learn dependencies grounded in physical principles and to produce physically consistent models with better chances of generalizing.The approach targets shortcomings of both data-only and theory-only models.
- Generalization: Scientific knowledge can prune physically inconsistent candidates, reducing model variance without likely affecting bias and focusing learning on generalizable, interpretable solutions.This reframes physical consistency as part of model performance alongside training accuracy and model complexity.
- Generalization: Performance ∝Accuracy + Simplicity + Consistency.The revised objective makes physical consistency an explicit component of TGDS model performance.
- Research themes: The paper describes five broad categories for combining scientific knowledge with data science, including model design, learning guidance, and augmentation of theory-based models.These approaches can range from strict enforcement of physical consistency to relaxed use of knowledge as priors or regularization.
3 THEORY-GUIDED DESIGN OF DATA SCIENCE MODELS
Theory-guided design incorporates domain knowledge into the model family itself, especially through response or loss specifications and architecture choices. These choices aim to produce solutions that are both easier to optimize and more physically meaningful.
- Design principles: Scientific knowledge can guide model-family selection by shaping response functions, loss functions, and model architectures.The paper emphasizes matching the data science model’s relationship form to domain understanding.
- Theory-guided specification: Synergistic response and loss functions can simplify optimization while remaining consistent with physical understanding and supporting generalizable solutions.The paper discusses this strategy in generalized linear models and artificial neural networks.
- Theory-guided specification: In generalized linear models, the expected target mean is linked to a weighted input combination through g(µ) = wT x + b, with w and b learned from data.Different link and probability-distribution choices produce different regression models.
- Theory-guided specification: For highly skewed extreme responses such as severe floods and droughts, a Gumbel distribution is more physically meaningful than the standard Gaussian assumption.The distribution should match the domain characteristics of the response variable.
- Theory-guided architectures: Domain knowledge can inform neural-network architectures by encoding biologically plausible rules, modular physical subprocesses, and theory-guided dependencies.Examples include view-invariant face representations, hydrological subprocess modules, and architectures capturing spatial or temporal structure.
4 THEORY-GUIDED LEARNING OF DATA SCIENCE MODELS
Theory-guided learning incorporates scientific knowledge into model initialization, probabilistic structure, constraints, and regularization. These approaches aim to guide models toward physically consistent, generalizable, and interpretable solutions across scientific applications.
- 4.1 Theory-guided Initialization: Theory-guided initialization uses domain knowledge or simulations to start iterative learning from physically consistent parameter values.Simulation-based pretraining is proposed for neural networks because scientific datasets are often small relative to the number of variables.
- 4.2 Theory-guided Probabilistic Models: Theory-guided probabilistic models incorporate scientific relationships through graphical-model structure or domain-informed priors.Such priors can reduce parameter variance in high-dimensional problems with few labeled examples, including electrophysiological imaging of the heart.
- 4.3 Theory-guided Constrained Optimization: Complex scientific constraints remain difficult for traditional constrained optimization when they involve PDEs or nonlinear variable transformations.The paper identifies domain-driven PDE approaches as a direction for incorporating memory effects and nonlinear, energy-conserving interactions into time-series regression.
- 4.3 Theory-guided Constrained Optimization: Theory-guided constrained optimization restricts model learning with scientific constraints, including elevation ordering, partial differential equations, and modified Euler–Lagrange conditions.In quantum chemistry, a self-consistent model must predict kinetic energy while also accurately estimating ground-state density; restricting density solutions to the training density manifold addresses inconsistent solutions.
- 4.3 Theory-guided Constrained Optimization: Elevation information can constrain water–land classification so that predicted labels obey physically viable ordering even with poorly labeled training data.If B is labeled as land, a higher-elevation location A should also be labeled as land rather than water.
- 4.4 Theory-guided Regularization: Theory-guided regularization replaces purely generic parameter penalties with domain-specific structure or physically consistent model subspaces.Group Lasso can preserve related attributes, such as climate variables measured at the same location, while inferred task structure can support multi-task learning when task composition is unknown.
ENCE OUTPUTS
TGDS can refine data-science outputs by enforcing explicit or inferred scientific constraints. Examples include physically consistent material discovery and water-body mapping using latent elevation structure.
- Explicit and implicit constraints: Scientific theories can refine data-science outputs to produce physically consistent results rather than merely reducing noise or missing-value effects.The paper presents this as a final-stage model-building strategy that incorporates domain knowledge into output refinement.
- Explicit and implicit constraints: Material discovery uses theory-guided refinement to identify novel crystal structures with desirable properties in an extremely large search space.The motivating properties include gas filtering and catalysis, while traditional approaches rely on computationally expensive ab initio calculations.
- Explicit and implicit constraints: For Lake Abhe, the workflow progresses from remote-sensing imagery and initial classification to inferred elevation contours and elevation-refined final maps.The figure presents the image, initial maps, inferred contours, and refined maps as stages (a) through (d).
- Explicit and implicit constraints: Surface-water mapping jointly infers hidden elevation constraints and uses them to refine classification outputs when explicit elevation data are unavailable.The approach estimates latent ordering among locations from long-term histories of imperfect water/land labels.
- Explicit and implicit constraints: Implicit domain constraints have also been used for urbanization and tree-plantation mapping through hidden Markov models representing land-cover transitions.
6 LEARNING HYBRID MODELS OF THEORY AND DATA SCIENCE
Hybrid TGDS models combine theory-based components with data-science components, assigning different parts of a scientific problem to each. Turbulence examples use machine learning to augment or correct Reynolds-averaged Navier–Stokes models.
- Hybrid model designs: Hybrid TGDS models combine theory-based and data-science components so that some aspects are handled by physical models and others by learned models.One design feeds theory-based outputs into data-science models; another predicts missing or inaccurate intermediate quantities for theory-based models.
- Turbulence modeling: Turbulence modeling remains difficult because exact simulations are computationally expensive and current RANS approximations inadequately represent separated, curved, or swirling flows.RANS models approximate the unknown Reynolds stress introduced by turbulent fluctuations.
- Turbulence modeling: A random forest can estimate a model discrepancy ΔτML that is added to Reynolds stress obtained from a RANS model.Because the discrepancy is learned independently, this approach does not change the form of the RANS approximation itself.
- Turbulence modeling: Field inversion and machine learning integrates neural-network terms into theory-based equations, exemplified by estimating effective viscosity with physical production, destruction, and transport terms.The paper identifies this framework as FIML and presents such coupling as a way to reduce discrepancies in complex scientific applications.
7 AUGMENTING THEORY-BASED MODELS USING DATA SCIENCE
TGDS can augment theory-based models by assimilating observations into model states and by calibrating high-dimensional parameters. These approaches aim to improve agreement with observations and make physical models more effective.
- Data assimilation: Data assimilation infers the most likely sequence of physical states by constraining each current state with previous states and current observations.Kalman filtering is a special case when linear transitions are modeled with Gaussian distributions; more complex dependencies can encode physical laws.
- Data assimilation: Data assimilation has been widely used in climate science and hydrology for integrating observations into dynamical theory-based models.
- Parameter calibration: Theory-based models often require parameter calibration to represent physical systems accurately, but exhaustive searches over parameter combinations become infeasible as dimensionality grows.The paper frames parameter calibration as a high-dimensional selection problem rather than a practical grid-search exercise.
- Parameter calibration: Multi-armed bandit techniques provide a promising direction for calibrating high-dimensional theory-based-model parameters using exploration, exploitation, and limited observations.These methods incrementally select parameter values while balancing exploration of choices against exploiting the highest-reward choice.
8 CONCLUSION
The paper establishes TGDS as a paradigm for systematically integrating scientific knowledge with data science, balancing strict physical consistency with relaxed knowledge use. It identifies practical benefits, broader research opportunities, and the need for deeper formalization to support future scientific discovery.
- TGDS integrates scientific knowledge with data science through approaches ranging from strict physical-consistency enforcement to relaxed priors and regularization.The taxonomy accommodates applications with varying strength of scientific understanding.
- Anchoring data science algorithms with scientific knowledge aims to improve model generalizability for complex problems with under-representative samples and produce scientifically interpretable models.Restricting learning to physically consistent models may also reduce computational cost.
- Future TGDS research should extend beyond the paper’s supervised-learning focus to uncertainty quantification, pattern mining, scientific workflows, and model simulation.The paper presents these as examples of non-exhaustive research directions for blending theory with data science.
- The paper is a first step toward deeper theoretical formalizations of TGDS, whose success depends on handling diverse forms, granularity, completeness, and uncertainty of scientific knowledge.The presented approaches are characterized as a stepping stone toward more deeply integrated theory-based and data-science methods.
- TGDS is positioned as a potential foundation for a fourth paradigm of scientific discovery in which data contributes throughout knowledge discovery.