Source-linked AI summary
Machine Learning and Materials Informatics: Recent Applications and Prospects
Rampi Ramprasad, Rohit Batra, Ghanshyam Pilania, Arun Mannodi-Kanakkithodi, Chiho Kim
TL;DR
Materials science needs data-driven methods for properties that are difficult or expensive to measure or compute, provided reliable data exists or can be generated. The article surveys materials-informatics strategies that fingerprint materials and learn surrogate property mappings, including multi-fidelity and inverse-design approaches. Reported applications include screening boron-containing perovskites near ∼1 GV/m and classifying Heusler compounds with a 94% true positive rate, while emphasizing uncertainty, out-of-domain prediction, and fingerprinting challenges.
Problem
Materials informatics addresses materials properties that are difficult, costly, or time-consuming to measure or compute when reliable data is available or can be generated for critical cases.
Method
The paper surveys data-driven strategies that fingerprint materials, learn surrogate mappings from fingerprints to properties, and extend these ideas to multi-fidelity learning and inverse design.
Results
Machine learning enabled screening of boron-containing perovskites predicted near ∼1 GV/m intrinsic breakdown strength and classified Heusler compounds with a true positive rate of 94%.
Takeaways & Limitations
Materials informatics can provide rapid surrogate predictions and support materials screening, while adaptive learning may improve models through systematic infusion of new data.
Takeaways & Limitations
Reliable past data is required, and major challenges remain in recognizing out-of-domain cases, quantifying uncertainty, and converting inverse-design fingerprints into physically meaningful materials.
Abstract
from arXiv · showhide
Propelled partly by the Materials Genome Initiative, and partly by the algorithmic developments and the resounding successes of data-driven efforts in other domains, informatics strategies are beginning to take shape within materials science. These approaches lead to surrogate machine learning models that enable rapid predictions based purely on past data rather than by direct experimentation or by computations/simulations in which fundamental equations are explicitly solved. Data-centric informatics methods are becoming useful to determine material properties that are hard to measure or compute using traditional methods--due to the cost, time or effort involved--but for which reliable data either already exists or can be generated for at least a subset of the critical cases. Predictions are typically interpolative, involving fingerprinting a material numerically first, and then following a mapping (established via a learning algorithm) between the fingerprint and the property of interest. Fingerprints may be of many types and scales, as dictated by the application domain and needs. Predictions may also be extrapolative--extending into new materials spaces--provided prediction uncertainties are properly taken into account. This article attempts to provide an overview of some of the recent successful data-driven "materials informatics" strategies undertaken in the last decade, and identifies some challenges the community is facing and those that should be overcome in the near future.
Overarching Perspectives
Materials informatics is emerging as an essential materials-research strategy, using reliable existing or targetedly generated data to build surrogate predictions while raising questions about problem suitability, uncertainty, and out-of-domain use.
- Historical examples include Hume-Rothery rules, Hall-Petch relationships, and group-contribution methods, which established empirical links between material descriptors and properties.
- Data-driven materials research has become an essential part of the materials research portfolio, supported by simulations and high-throughput experiments or computations that generate targeted critical data.
- Materials informatics learns from past reliable data to identify previously unknown correlations and make rapid property predictions without directly solving governing equations.
- Applying models outside the domain of prior data is dangerous, making recognition of out-of-domain cases and prediction-uncertainty quantification central challenges.
- The article reviews successful data-driven materials strategies from the last decade and identifies challenges the community should overcome.
Elements of Machine Learning (Within Materials Science)
Materials-science machine learning begins with reliable materials-property data, converts materials into numerical fingerprints, and learns mappings from those fingerprints to target properties using validated surrogate models.
- Reliable, curated past data or a controlled effort to create it is a prerequisite for applying machine learning to a materials problem.
- A materials-property dataset defines inputs as materials and targets as measured or computed properties, enabling prediction of a new material’s property.
- Fingerprinting reduces each material to a numerical representation, a domain-expertise-intensive step that makes quantitative prediction possible.
- Learning establishes a numerical mapping from fingerprints to continuous or discrete target properties using algorithms such as regression, trees, or neural networks.
- Cross-validation and testing on unseen data are essential for checking generalization and avoiding overfitting.
- The broader machine-learning process includes organized dataset creation, fingerprinting, learning, and progressive targeted data infusion for adaptive improvement.
Hierarchy of Fingerprints or Descriptors
Fingerprint granularity should match the problem’s goals and accuracy requirements: finer representations can improve accuracy but demand more labor, data, and computation, whereas coarser ones support rapid screening.
- Fingerprint design depends on the problem and prediction goals, with granularity varying according to required accuracy and the desired level of understanding.
- Finer fingerprints generally offer greater expected accuracy but require more labor and data and provide less conceptual learning.
- Coarser fingerprints are generally appropriate for rapid initial screening, while finer representations suit accuracy-critical predictions.
- Material fingerprints should remain invariant to rigid translations and rotations so equivalent physical configurations receive equivalent representations.
Examples of learning based on gross-level property-based descriptors
Gross-level descriptors support surrogate models for complex materials properties, including breakdown strength, crystal-structure preference, and other thermophysical and mechanical quantities. These approaches combine physically meaningful primary features with nonlinear feature generation, selection, and predictive modeling.
- Hume-Rothery and Hall-Petch relationships show how coarse descriptors can connect composition or grain size with materials behavior.Modern machine learning extends this knowledge-discovery strategy to multivariate and highly nonlinear dependencies.
- LASSO-based workflows generate large nonlinear descriptor spaces, select critical features, and build surrogate models for intrinsic electrical breakdown fields.The workflow includes cross-validation and testing after feature generation and down-selection.
- The breakdown-field model screened perovskites and identified boron-containing compounds predicted to have intrinsic breakdown strengths of ∼1 GV/m, later confirmed by first-principles computations.The benchmark began with 82 binary octet insulators, while four new cases were outside the original dataset.
- Descriptor-based classification predicted preferred crystal structures, including an 86% correct-structure probability for 2,105 experimentally known AxBy examples.The model searched 1.7 × 10^5 nonlinear descriptors and reduced them to three optimal descriptors.
- A related proof of concept predicted 12 novel gallide candidates, which were synthesized and confirmed as Heusler compounds.The candidates had formulae MRu2Ga and RuM2Ga, with M = Ti, V, Cr, Mn, Fe, or Co.
- Gross-level descriptors have also supported surrogate models for band gaps, formation enthalpies, free energies, defect energetics, melting temperatures, mechanical properties, thermal conductivity, and catalytic activity.
Examples of learning based on molecular fragment-level descriptors
Molecular fragment-level fingerprints represent materials through constituent building blocks and enable learning polymer and inorganic-material properties. Polymer studies use chemically allowed repeat-unit combinations, DFT-computed dielectric properties, and kernel-based prediction schemes.
- Fragment-level descriptors encode finer structural detail than gross properties by representing materials through molecular or chemical building blocks.Their origins include cheminformatics, QSAR/QSPR, and polymer group-contribution methods.
- Hundreds of polymers assembled from seven chemically allowed basic units were evaluated with DFT for dielectric constant and band gap prediction.The study included van der Waals interactions and fingerprinted the polymers for machine-learning modeling.
- Kernel ridge regression uses distances in fingerprint space between a new polymer and training examples to predict its property.The workflow is illustrated for surrogate predictions of key dielectric polymer properties and implemented in the Polymer Genome application.
- Polymer properties that are difficult to compute or measure, including band gap, dielectric constant, glass transition temperature, and dielectric loss, can be learned and predicted.
- Fragment-based representations extend beyond polymers to predicting compositions and likely crystal structures of AxByOz ternary oxides.The probabilistic model combines crystal-structure type with elemental composition information.
Examples of learning based on sub-Angstrom-level descriptors
Sub-Angstrom-level fingerprints capture detailed atomic configurations for high-fidelity property and structure prediction. They underpin machine-learning force fields, chemical-accuracy targets, direct force learning, and fine-scale structural classification, while requiring demanding invariance and smoothness properties.
- Fine-level fingerprints encode precise atomic configurations for learning potential energies, structural phases, motifs, and atomic forces.Such models can support accelerated atomistic computation and efficient on-the-fly characterization.
- Chemical accuracy means errors below 1 kcal/mol for potential energies and reaction enthalpies, or below 0.05 eV/˚A for atomic forces.These thresholds are important for reliable molecular dynamics and precise structural-phase identification.
- Machine-learning surrogates trained on atomic configuration-to-property data can be several orders of magnitude faster than DFT while retaining quantum-mechanical and chemical accuracy.They also address the limited transferability of traditional semi-empirical force fields when adequate reference DFT data are available.
- Fine-level fingerprinting schemes must be invariant to translations, rotations, and exchange of like atoms, while remaining continuous and differentiable under small positional changes.Proposed approaches include symmetry functions, bispectra, Coulomb matrices, and SOAP.
- Machine-learning force fields map fingerprints to materials properties using neural networks, kernel ridge regression, or Gaussian process regression.Behler-type schemes and Gaussian approximation potentials have demonstrated chemical accuracy, versatility, and efficiency.
- Direct atomic-force learning determines total potential energy by integrating forces along a reaction coordinate or molecular-dynamics trajectory.The approach assigns forces uniquely to individual atoms.
- Fine-level fingerprints also support structure refinement, phase-diagram construction from XRD spectra, and classification of atomic environments.Examples include Bayesian/MCMC refinement and SOAP-based dimensionality reduction for silicon.
Critical steps going forward
Materials informatics advances by combining uncertainty-aware adaptive design with algorithms that handle complex outputs, multiple fidelity levels, and inverse design. Progress depends on matching problems to reliable data and respecting the limits of extrapolation and physical realizability.
- Adaptive design and uncertainty: Predicted uncertainty can reveal whether a new case lies inside or outside the training-data domain and distinguish interpolation from extrapolation.Uncertainty quantification also supports decisions about which new materials to add during iterative learning.
- Adaptive design and uncertainty: Adaptive design improves models by balancing exploration of new material systems against exploitation of currently promising candidates.The goal is to select the next material or question that improves the model or target-material search.
- Advanced learning targets: Vectorial targets such as electronic or vibrational densities of states are better learned as complete functions rather than as independent scalar values.The approach seeks to capture the target function simultaneously across energy or frequency.
- Multi-fidelity learning: Multi-fidelity learning combines properties computed at different cost and accuracy levels rather than requiring low-fidelity values for every material.This addresses combinatorial searches across large chemical and configurational spaces, where exhaustive low-fidelity computation can be demanding.
- Inverse design: Inverse design seeks materials with desired properties, but converting predicted fingerprints into physically and chemically meaningful materials remains a major hurdle.Proposed strategies constrain fingerprints to realizable materials or iteratively search material populations with genetic algorithms or simulated annealing.
- When to use machine learning: Machine learning requires reliable past data and is most appropriate for properties costly to measure or compute, complex phenomena, or settings with unknown governing equations.Problem selection should precede adoption of machine-learning methods.