Source-linked AI summary
Ab initio machine learning in chemical compound space
Bing Huang, O. Anatole von Lilienfeld
TL;DR
Chemical compound space is too large for broad first-principles virtual sampling, while existing data and model choices face computational, selection, and representation constraints. The review synthesizes quantum-mechanics-based machine-learning approaches that use quantum data and physics-informed models to accelerate property prediction across CCS. It concludes that QML can support rapid surrogate predictions, but general rational computational discovery and design has not yet been achieved.
Problem
The colossal size of chemical compound space makes first-principles sampling prohibitive, while random training-set selection can introduce bias and redundant or irrelevant instances.
Method
The review surveys QML representations, regressors, and training-set strategies built from quantum-mechanical data and designed for property prediction across chemical compound space.
Results
QML developments provide surrogate quantum-property predictions and can reduce costly high-level quantum data requirements, including an approximately 10-fold reduction for chemical accuracy on out-of-sample QM7b atomization energies.
Takeaways & Limitations
QML combines statistical-model efficiency with an ab initio perspective, but the general goal of rational computational discovery and design of compounds with desired properties remains unmet.
Takeaways & Limitations
Published high-quality datasets remain limited to very few or small molecules, typically no more than three heavy atoms, and the review cannot guarantee completeness of its rapidly growing outlook.
Abstract
from arXiv · showhide
Chemical compound space (CCS), the set of all theoretically conceivable combinations of chemical elements and (meta-)stable geometries that make up matter, is colossal. The first principles based virtual sampling of this space, for example in search of novel molecules or materials which exhibit desirable properties, is therefore prohibitive for all but the smallest sub-sets and simplest properties. We review studies aimed at tackling this challenge using modern machine learning techniques based on (i) synthetic data, typically generated using quantum mechanics based methods, and (ii) model architectures inspired by quantum mechanics. Such Quantum mechanics based Machine Learning (QML) approaches combine the numerical efficiency of statistical surrogate models with an {\em ab initio} view on matter. They rigorously reflect the underlying physics in order to reach universality and transferability across CCS. While state-of-the-art approximations to quantum problems impose severe computational bottlenecks, recent QML based developments indicate the possibility of substantial acceleration without sacrificing the predictive power of quantum mechanics.
I. INTRODUCTION
The review frames chemical compound space as an enormous set of stable compounds whose exploration is computationally prohibitive, motivating quantum machine learning approaches. It introduces QML as statistical learning of quantum properties across CCS and situates it alongside potential-energy-surface modeling and flexible data-driven alternatives to fixed force fields.
- Review scope: The review surveys machine-learning models that train and predict quantum properties throughout chemical compound space.It distinguishes this use of statistical learning from machine learning applied to quantum computing.
- Chemical compound space: Chemical compound space comprises theoretically conceivable, metastable compounds defined by locally stable atomic configurations separated by barriers against spontaneous reactions.Its size grows explosively with atom count through combinatorial stoichiometries and energetic diversity.
- Related modeling: Related potential-energy-surface studies usually model a single system and compute energies or forces from scratch, without exploiting relationships across constitutional and compositional isomers.The review treats these studies as complementary to CCS-wide QML.
- Modeling limitations: Flexible data-driven fitting models are motivated partly by force fields whose fixed functional forms are difficult to improve with more data and can fail catastrophically in some regimes.Earlier Shepard interpolation work already addressed data selection and the trade-off between accuracy and training cost.
C. Navigating CCS from first principles
First-principles navigation of chemical compound space builds on quantum-property datasets and representations that connect nuclear composition and structure to property changes. The reviewed approaches include quantum high-throughput design, locality-based QML, and alchemical methods that vary nuclear charges or electron number.
- Motivation: Improved hardware and quantum approximations have produced sizable quantum-mechanical datasets that support rapid property estimates for new compounds with statistical surrogate models.The research question is how properties trend across chemical compound space.
- Materials design: First-principles high-throughput computational design has become an important success story, with machine learning used to discover new ternary materials databases.The cited machine-learning efforts date to seminal work in 2010.
- Locality: QML maps molecular distance or similarity to property differences under a locality assumption, requiring test-set nuclear types to be represented in training.Predictive performance depends on similarity between local and global entities.
- Alchemical methods: Alchemical approaches extend first-principles exploration by treating nuclear charges, proton number, or electron number as variables in electronic-structure models.The literature includes variable-charge density-functional formulations and later applications to redox potentials, derivatives, and exchange-correlation potentials.
- Applications: Quantum alchemical changes have been applied to stability, thermodynamic integration, mixtures, reactivity, chemical-space exploration, binding, adsorption, and electronic properties.Related extensions also address kinetic isotope effects beyond the Born-Oppenheimer approximation.
II. HEURISTIC APPROACHES
Heuristic approaches use simplified chemical representations and statistical relationships to model properties, often effectively within restricted domains. Their usefulness is balanced by limited transferability, representation gaps, and dependence on assumptions that may not reflect quantum mechanics.
- QSPR: Conventional QSPR methods rely predominantly on heuristic assumptions about the forward problem and are therefore limited to particular applicability domains.Their implicit bias is associated with lacking a basis in the underlying physics.
- QSPR: Heuristic QSPR can still reveal qualitative trends and sometimes predict accurately for specific property sub-domains and systems.The passage explicitly distinguishes this utility from direct reliance on quantum-mechanical laws.
- Approach families: The review organizes heuristic literature into low-dimensional correlations, coarse molecular representations, and representations based on molecular properties.These perspectives are presented largely in chronological order.
- Coarse models: Simple one- or few-variable models act as coarse-grained schemes that can capture essential physics in specific chemical sub-domains but lack quantum-mechanical transferability.The review notes that such models may remain useful for hard problems such as intensive properties and multireference character.
- Stoichiometry: For fixed structural patterns, stoichiometry can uniquely represent systems; an exhaustive QML scan of 2 million elpasolite crystals predicted nearly 90 favorable crystals later added to the Materials Project.The stoichiometric representation reached explicit geometry-based many-body representations at larger training-set sizes.
- Connectivity: When systems lack a common structural skeleton, stoichiometry alone is insufficient and bonding connectivity or conformation may also be required.Graph-based representations remain limited for noncovalent interactions and reactions involving graph transformations.
D. Coarse-grained
Coarse-grained representations address the rising cost of QML for larger systems by grouping nearby atoms into superatoms or beads. Descriptor-based models offer broad applicability but can lose predictive power when representations are non-unique, while supervised QML learns property mappings from structure-property data.
- Coarse-grained: As system size grows, QML training and prediction become more expensive, although they scale more favorably than typical quantum-chemistry methods.Beyond certain size thresholds, direct treatment may become very demanding or impossible.
- Coarse-grained: Coarse-grained models represent groups of nearby atoms as superatoms or beads to make large molecular and soft-matter systems tractable.Coarse-grained QML is presented as a systematic alternative to coarse-grained force fields when accuracy control is needed.
- Property-based: Descriptor representations select calculated or measured atomic or molecular properties as features and commonly pair them with nonlinear regressors such as neural networks.The selected properties are intended to describe the target property and must be relatively easy to obtain.
- Property-based: Descriptor-based predictive power is limited by construction when the chosen representation lacks uniqueness, despite potential applicability across system sizes and compositions.This limitation follows from the representation rather than the choice of regressor alone.
- QML methodology: Supervised QML learns a quantum-property mapping from compounds paired with reference properties and applies it to new, out-of-sample compounds.The reference data may be calculated or measured, and some inferred properties need not be observables.
A. Regressor
QML regressors map chemical compound space into similarity-based representations and learn property predictions from quantum-mechanical reference data. Learning efficiency depends on the regressor, metric, representation, data quality, and effective dimensionality.
- Regressor: Kernel ridge regression and related models fit generic basis-function expansions to pre-calculated quantum properties, with kernel methods offering comparatively light-weight training for scarce data.Deep neural networks generally require much larger data sets and more optimization effort, whereas kernel methods can be faster to train.
- Regressor: Correctly regularized models trained on noise-free data can interpolate reference properties and make statistically meaningful predictions for out-of-sample compounds.Converged cross-validation and hyperparameterization are required for this behavior.
- Learning curves: E ∝ a/N^b describes the leading-order decay of out-of-sample prediction error with training-set size for GPR/KRR and neural-network models.On log-log scales, learning curves follow log E = log a − b log N and support systematic comparisons of model efficiency.
- Learning curves: Representation choices can cause learning to cease or shift learning-curve offsets when information is incomplete, non-unique, noisy, or physically mismatched.Off-diagonal elements following an R^-6 power law achieved lower offsets than Coulomb-like decay, while linearly or quadratically growing elements increased offsets.
- Learning curves: Expert-informed removal of irrelevant information can lower effective dimensionality and produce substantially steeper learning curves than ordinary unique representations.The review identifies rational training-set sampling as a strategy for exploiting this compactness.
C. Loss-functions
QML performance depends strongly on how representations encode physical information and on whether the loss function matches the prediction target. Derivative information helps some potential-energy-surface tasks, but its value varies across chemical properties and prediction settings.
- Loss-functions: Derivative information can dramatically improve potential-energy-surface fitting, but its benefit depends on the target: gradients negligibly improve CCS atomization-energy learning while increasing kernel-basis requirements.For predicting gradients throughout CCS, training with energies alone offers no advantage over training with forces.
- Representations: Representation uniqueness is necessary to avoid spurious noise from neglected degrees of freedom and to reduce prediction errors below conformational-property variance.Covalent-connectivity graphs cannot capture conformational degrees of freedom, imposing a corresponding error floor.
- Representations: Representations should ideally combine compactness, computational efficiency, symmetries, invariances, and physical meaning because they define the basis functions and learning-curve shape.Encoding target-property invariances typically decreases the learning-curve offset.
- Representations: A single representation and kernel can model all quantum-mechanical properties, with property-specific regression coefficients determined by the reference property vector.This property-independent setup contrasts with conventional QSAR/QSPR models whose regressors typically depend strongly on the target property.
- Representations: Predictive accuracy varies widely with representation and regressor choice, as illustrated by successive improvements in QM9 atomization-energy models.The review presents these benchmarks as evidence that model design materially affects QML accuracy.
- Representations: Including increasingly more physical information in a representation systematically improves learning curves, while less physical representations worsen them.These developments were benchmarked on atomization energies of small organic molecules in QM9.
B. Continuous
Continuous and distribution-based representations address atom-indexing problems in QML while incorporating structural and alchemical information. FCHL18 models demonstrated transfer to molecules containing elements absent from training.
- Continuous representations: Sorting atoms can enforce indexing invariance in discrete representations but may introduce derivative discontinuities that harm force predictions.The problem arises because sorting is artificial and can make the representation non-smooth.
- Continuous representations: Continuous or distribution-based representations overcome sorting by integrating atom-index-dependent distances, angles, dihedrals, or nuclear charges through smeared projections.These representations are closely related to many-body or cluster expansions.
- FCHL18: FCHL18 adds alchemical degrees of freedom to a structural distribution-based many-body representation.The name denotes the FCHL representation introduced in 2018.
- FCHL18: FCHL18 encodes two-body distance distributions with r^-4 scaling and three-body angular distributions with r^-2 scaling.Four-body terms were tested but had negligible impact on learning curves.
- Transfer across chemical space: FCHL18-based QML models accurately inferred properties for systems containing chemical elements absent from training.For H_nY∼X molecules, semi-quantitative covalent-bond potential curves were predicted after training on DFT curves for molecules containing neither X nor Y.
A. ML models of parameters
QML extends existing computational models through learned parameters, corrections, and multi-fidelity schemes. These approaches can reduce high-level data demands, but their effectiveness depends on baseline quality and the relationship between theory levels.
- ∆-ML: ∆-ML learns label corrections across chemical compound space and was shown to improve systematically with training-data size.The approach modeled energy and geometry differences across multiple theory levels, including PM7 through CCSD(T).
- ∆-ML: ∆-ML also corrected complex properties in data-scarce settings, including van der Waals interactions and higher-order catalyst-activity estimates.One example used fewer than one hundred training instances for DFT corrections.
- ∆-ML: Learning curves show a constant offset reduction for noise-free data and functional representations, whereas corrections of coarse baselines can yield vanishing improvement.The latter behavior was reported for PM7 and Hammett’s relation.
- Multi-fidelity ML: Multi-fidelity ML combines a broad low-level training set with a smaller high-level set by learning energy differences between theory levels.Recursive KRR trains a baseline on low-level data and a second model on high-minus-low-level energy differences.
- Multi-fidelity ML: Recursive multi-fidelity KRR predicts high-level properties by summing inferred contributions across adjacent theory levels.For levels 0 through L, the training sets are nested as S0 ⊂ S1 ⊂ ··· ⊂ SL.
- Multi-fidelity ML: MF-GPR differs from MF-KRR by modeling each level as a Gaussian process and providing prediction-variance estimates.MF-KRR converges toward conventional KRR at the highest level when differences between training sets vanish.
D. Multi-level-grid-combination
Multi-level-grid-combination (MLGC) uses sparse-grid ideas to combine quantum calculations across correlation levels, basis sets, and training-set sizes. It reduces the costly highest-level data needed for chemically accurate predictions, while introducing challenges in defining distances between abstract theory variables and selecting suitable transfer-learning models.
- Multi-level-grid-combination: MLGC expresses system properties across abstract variables such as electron-correlation level, basis set, and training-set size.The formulation extends from E = E(xC, xB) to E = E(xC, xB, xN), where xN denotes training-set size.
- Multi-level-grid-combination: Sparse-grid interpolation and extrapolation combine calculations at different correlation levels and basis sets, including unsampled high-accuracy regions.The approach targets combinations with large xC and xB that were not directly sampled.
- Multi-level-grid-combination: The abstract-variable formulation lacks a natural quantitative distance between correlation methods, complicating sparse-grid weighting.For example, the relative distance between HF, MP2, and CCSD(T) along the correlation dimension is not quantitatively established.
- Multi-level-grid-combination: ∼10 fold reduction in CCSD(T)/cc-pVDZ data enabled chemical-accuracy atomization-energy predictions for out-of-sample QM7b molecules.The comparison is against a traditional single-level ML model.
- Transfer learning: Transfer learning reuses knowledge from a related base task, such as DFT-energy prediction, to initialize a target task such as CCSD(T)-energy prediction.Inductive transfer learning keeps source and target domains the same while changing the task.
- Transfer learning: Transfer-learning performance depends on choosing an appropriate base model and dataset, and negative transfer can degrade the target task.Freezing is suggested when target labels are scarce, whereas fine-tuning is suggested when labels are more available.
VI. TRAINING SET SELECTION
Training-set selection is central because QML knowledge is encoded in the training data, yet CCS grows combinatorially and is heterogeneous. The review frames representative, suitable, and system-specific selection as open problems and notes that random selection inevitably suffers selection bias.
- Training-set selection: Training-set selection is fundamentally important because the data implicitly encode the knowledge required for confident QML predictions.The issue becomes more difficult as system size, elemental diversity, and CCS complexity increase.
- Open questions: Three open questions concern extracting representative nonredundant subsets, quantifying suitability for a query, and selecting data for a specific system.These are labeled Q1, Q2, and Q3 in the review.
- Selection strategies: Random selection is broadly applicable but inevitably suffers from selection bias when constructing training sets.The nonlinear effect of individual training instances on model parameters makes systematic selection challenging.
- Sources of bias: CCS bias arises from the curse of dimensionality and inhomogeneity, with combinatorial growth driven by system size and compositional diversity.The passage identifies these as two distinct components of the bias problem.
- Existing approaches: Existing approaches often assume a pre-existing or easily generated dataset, including genetic algorithms and active-learning methods.These methods respectively optimize labeled subsets or select representative unlabeled configurations before labeling.
A. Genetic algorithm
Genetic algorithms select QML training sets by iteratively evaluating candidate subsets and applying selection, crossover, and mutation. Optimized sets improve generalization over random sampling, but the fitness evaluation generally requires labeled data; active learning instead targets informative unlabeled points through uncertainty-based criteria.
- Genetic algorithm: Genetic algorithms generate candidate subsets, train QML models, evaluate test error as fitness, and evolve populations through selection, crossover, and mutation.Mutation can replace functional groups, such as -NH2 with -CH3, to promote diversity.
- Genetic algorithm: The converged subset is intended to represent typical atomic environments and improve model performance relative to randomly drawn training sets.The optimized set’s usefulness is assessed by generalization to molecules absent from the original pool.
- Genetic algorithm: Improved generalizability was observed for PubChem molecules compared with random sampling.The reported comparison evaluates models trained on the optimized subset against randomly selected training sets.
- Genetic algorithm: Most genetic-algorithm implementations require labeled data to evaluate fitness, reducing QML-model cost but not the total need for training data.This is identified as the main limitation of the approach.
- Active learning: Active learning can select unlabeled points before costly labeling, with variance-reduction strategies choosing points expected to reduce predictive uncertainty.Variance estimation is distinct from mean-error estimation.
- Active learning: D-optimality exploits lower-dimensional feature representations and linear local atomistic potentials to select informative molecular subsets.The framework can incorporate forces by differentiating basis functions with respect to Cartesian coordinates.
- Active learning: Gaussian-process regression provides a direct variance estimate for deciding whether a new point is likely to improve the model.Large variance indicates potential value relative to a user-defined tolerance, whereas small variance suggests limited benefit.
- Active learning: Neural-network uncertainty estimates tend to be overconfident because standard models often output a single prediction rather than a predictive distribution.The review contrasts this with the uncertainty estimates available from methods such as GPR.
C. AMON based QML
AMON-based QML selects query-specific molecular fragments to represent local chemistries, reducing the training data needed for predictions across chemical compound space. The approach combines graph-based fragment enumeration with geometry relaxation and quantum-chemistry validation.
- AMON selection: AMONs select an optimal query-specific training set on-the-fly instead of relying on randomly sampled starting data.The approach exploits locality to reconstruct extensive properties such as ground-state energy.
- AMON selection: The procedure constructs a query connectivity graph, enumerates increasing-size subgraphs, checks connectivity, isomorphism, and valence saturation, then relaxes candidate geometries.Dissociated or connectivity-altered fragments are discarded or rechecked.
- AMON selection: The resulting AMON set is considered representative of all local chemistries in the query molecule.The set is built by continuing through the enumerated subgraphs until they are exhausted.
- Illustrative example: For 2-(furan-2-yl)propan-2-ol, the procedure produced only 30 AMONs that collectively represent the target’s complete set of local atomic environments.These fragments may support accurate extrapolation to the target and other molecules sharing the same AMONs after fragmentation.
- Performance and scope: AMON models achieved improved learning-curve slopes and offsets for thousands of molecules with training sets of only ∼50 on average, versus twenty times larger random-sampling sets.The graph-based approach is best suited to compositional space; extending it to configurational space or graph-free systems remains nontrivial.
- Property prediction: QML studies pair molecular representations with labels spanning properties such as atomization energies, dipole moments, and boiling points.The review also discusses multi-property neural networks that encode correlations among quantum properties.
A. Atomic
Atomic QML properties generally benefit from locality and have been studied across molecular, solid-state, and biomolecular settings. Electronic and excited-state properties remain comparatively sparse across chemical compound space, although newer models broaden this coverage.
- Atomic properties: Atomic properties are relatively easy to learn because they benefit strongly from the locality of an atom within a molecule.Examples include core-level excitations, forces, NMR shielding constants, charges, dipole moments, and polarizabilities.
- Atomic properties: NMR-shift QML expanded from molecules to solids, solvated proteins, couplings, and broader benchmark studies between 2015 and 2020.The cited work includes molecular, solid-state, and solvated-protein applications.
- Electronic properties: QML models for electronic properties across CCS have remained rather sparse, including transport, impurity, dynamical-mean-field, and excitation-energy applications.The review distinguishes these CCS-wide models from recent studies of excited-state dynamics in given systems.
- Electronic properties: Recent developments include models for nonadiabatic excited-state dynamics and SchNarc, which combines SchNet with the SHARC surface-hopping code.Initial CCS studies with SchNarc involved small sets of small molecules.
- Electronic properties: Deep neural networks have been proposed for electron affinities and ionization potentials, while symmetry-conserving networks target electronic and vibrational spectra.These developments were reported in 2020.
C. Inter-molecular
Inter-molecular QML covers binding, reaction, catalytic, and related properties, but reaction barriers are harder because they involve off-equilibrium configurations and undersampled training spaces. Progress depends on datasets spanning increasingly diverse structures, geometries, and quantum properties.
- Inter-molecular energetics: Inter-molecular energetics include binding energies, reaction energies, and reaction barriers, with the latter involving significant structural reconstruction.The review focuses primarily on energetic properties in this category.
- Inter-molecular energetics: QML has been applied to formation energies in inorganic materials, chemical bonds in molecules, transition-metal complexes, surface reconstructions, and organic molecules.GPR and KRR provide a unified approach across several of these applications.
- Reaction properties: Reaction-barrier prediction is more difficult because off-equilibrium configurations are involved and the training space is undersampled.These conditions constrain accurate QML modeling of reaction-related properties.
- Catalysis: Catalysis applications include estimating adsorbate free energies for Pourbaix diagrams and searching for active-site motifs for CO2 reduction.Both GPR and neural-network models have been used for these tasks.
- Datasets: General QML applications require pre-existing training data, although generating data only when needed would minimize quantum-mechanical computations.The review surveys datasets designed to encode quantum information throughout CCS.
- Datasets: AMON-based AGZ7 fragments larger GDB17 and ZINC molecules into building blocks of no more than seven heavy atoms to alleviate the curse of dimensionality.Most datasets otherwise focus on equilibrium geometries, limiting direct coverage of dynamics and reactivity.
B. PubChem amd ZINC
PubChem and ZINC complement synthetic GDB-derived datasets with large databases oriented toward experimentally documented chemistry and drug design. Their associated quantum datasets broaden coverage, but practical QML still faces gaps in stability, synthesizability, excited states, reactions, and non-covalent interactions.
- PubChem and ZINC: GDB compounds are mostly virtual graph-enumeration products whose thermodynamic stability and synthesizability have not been established.Such local environments may be infeasible within complete molecular frameworks.
- PubChem and ZINC: PubChem contains over 111 million unique chemical-structure records from hundreds of sources, while PubChemQC supplies quantum data for approximately four million molecules.PubChemQC includes ground-state geometries and properties plus low-lying excited states.
- PubChem and ZINC: ZINC focuses on biochemical and drug-design applications, while AZ7 provides optimized quantum properties for ZINC AMONs with up to seven non-hydrogen atoms.AZ7 can serve as a fragment scaffold covering local ZINC chemistries.
- Additional databases: OE62 contains 61,489 molecules with energies and orbital eigenvalues computed at multiple theory levels, including vacuum and implicit-solvent settings.The dataset is based on the Cambridge Structural Database.
- Additional databases: Reaction datasets remain sparse, with QMrxn covering selected SN2 and E2 substituents and another dataset providing reactant, product, and transition-state properties.These datasets target narrow reaction subsets rather than broad reaction coverage.
- Additional databases: QMspin adds singlet and triplet carbene structures with optimized geometries and singlet-triplet spin gaps, while QMt records quantum-chemistry computational timings.QMt covers single-point energies, geometry optimizations, and transition-state searches for thousands of QM9 molecules.
- Additional databases: Public non-covalent-interaction datasets cover 3,700 distinct interacting molecule-pair types, while artificial datasets can provide soft labels for CCS benchmarking.These resources address interactions and chemically constrained synthetic spaces not covered by standard equilibrium datasets.
D. Transition metals
Transition-metal and periodic-system chemical spaces are difficult to explore because of complex compositions, geometries, and electronic structures. The review surveys representative datasets, computational properties, software, and workflows supporting QML studies across these domains.
- Transition-metal complexes: Transition-metal complexes are challenging to explore because their complicated electronic structures increase computational cost, with current efforts often using relatively low-level DFTB or DFT methods.The section identifies tmQM and related datasets as examples of this constrained exploration.
- Transition-metal complexes: tmQM contains 86,665 mononuclear complexes spanning diverse ligands and 30 transition metals, with geometries and electronic properties computed using DFTB and TPSSh-D3BJ/def2-SVP.Reported properties include orbital energies, dipole moments, and atomic charges.
- Transition-metal complexes: Kulik-group datasets cover multiple first-row transition metals and ligands ranging from weak-field chloride to strong-field carbonyl, alongside intermediate-field sulfur-, nitrogen-, and oxygen-containing ligands.
- Periodic systems: QMOF provides computed energies, band gaps, charge densities, and densities of states for more than 14,000 experimentally synthesized metal-organic frameworks spanning nearly the periodic table.
- Periodic systems: Solid and surface systems remain difficult because compositional, structural, and electronic diversity requires large-scale DFT calculations for structures, response properties, and thermodynamic quantities.The listed quantities include densities of states, band gaps, elastic tensors, bulk moduli, vibrational spectra, free energies, specific heats, and entropies.
- Datasets and software: The reviewed ecosystem includes large materials and catalyst datasets, quantum-code acceleration tools, standalone representation packages, and workflow platforms for constructing, managing, sharing, and reproducing computational data.Examples include AFlow, OQMD, Materials Project, Open Catalyst Project, QMLcode, PLUMED, and AiiDA.
X. COMPOUND DISCOVERY
Compound discovery can use sequential screening accelerated by QML or inverse design targeting desired properties. The review concludes that broad rational discovery remains unresolved, despite progress and promising directions such as alchemical QML.
- Discovery strategies: Brute-force discovery screens candidate compounds by solving Schrödinger equations and ranking computed properties, with QML potentially replacing repeated ab initio calculations after suitable training.
- Discovery strategies: Inverse design searches for compounds matching target properties and can use gradient-based global optimization to explore broader chemical subspaces.
- Discovery strategies: Current gradient-based inverse-design studies mostly use SMILES or molecular-graph features rather than 3D geometries, placing them in the QSPR regime because the representation-to-property mapping is not unique.
- Open challenges: Ground-state energies and forces for novel, distorted, charged, multireference, or long-range-interacting systems remain unresolved at simultaneously high efficiency and accuracy.The review identifies energy ranking of competing real-material structures as an important test for subsequent QML applications.
- Future directions: Alchemical perturbation theory could extend QML by learning energies and energy gradients with respect to nuclear charges while incorporating periodic-table composition alongside structural degrees of freedom.Preliminary FCHL results inferred properties of elements absent from training.
- Open challenges: Major unresolved issues include inadequate dataset-completeness theory, limited models for intensive properties, scarce experimental-quality datasets for medium-sized molecules, incomplete coverage of emerging directions, and the unachieved goal of general rational discovery.