Source-linked AI summary

Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems

Xuan Zhang, Limei Wang, Jacob Helwig, Youzhi Luo, Cong Fu, Yaochen Xie, Meng Liu, Yuchao Lin, Zhao Xu, Keqiang Yan, Keir Adams, Maurice Weiler, Xiner Li, Tianfan Fu, Yucheng Wang, Alex Strasser, Haiyang Yu, YuQing Xie, Xiang Fu, Shenglong Xu, Yi Liu, Yuanqi Du, Alexandra Saxton, Hongyi Ling, Hannah Lawrence, Hannes Stärk, Shurui Gui, Carl Edwards, Nicholas Gao, Adriana Ladera, Tailin Wu, Elyssa F. Hofgard, Aria Mansouri Tehrani, Rui Wang, Ameya Daigavane, Montgomery Bohde, Jerry Kurtin, Qian Huang, Tuong Phung, Minkai Xu, Chaitanya K. Joshi, Simon V. Mathis, Kamyar Azizzadenesheli, Ada Fang, Alán Aspuru-Guzik, Erik Bekkers, Michael Bronstein, Marinka Zitnik, Anima Anandkumar, Stefano Ermon, Pietro Liò, Rose Yu, Stephan Günnemann, Jure Leskovec, Heng Ji, Jimeng Sun, Regina Barzilay, Tommi Jaakkola, Connor W. Coley, Xiaoning Qian, Xiaofeng Qian, Tess Smidt, Shuiwang Ji

arXiv:2307.08423v6cs.LGphysics.comp-ph

TL;DR

AI for science is a large, interdisciplinary field lacking a unified technical treatment. This paper surveys AI for quantum, atomistic, and continuum systems, emphasizing shared challenges and symmetry-aware methods, and concludes that the area represents an emerging paradigm for scientific discovery.

  • Problem

    AI for science is an enormous, emerging, interdisciplinary field requiring unified treatment, while existing coverage remains necessarily selective.

  • Method

    The paper synthesizes AI methods and challenges across quantum, atomistic, and continuum systems, emphasizing equivariance to physical symmetries.

  • Results

    The paper presents AI for science as an emerging paradigm for scientific discovery spanning complex equations and expensive experimental observations.

  • Takeaways & Limitations

    The synthesis frames AI for science as a major interdisciplinary research area connecting computational methods with natural-science understanding.

  • Takeaways & Limitations

    The review is not comprehensive or conclusive and covers only selected AI-for-science areas, with relevant literature potentially omitted.

Abstract

from arXiv · show

Advances in artificial intelligence (AI) are fueling a new paradigm of discoveries in natural sciences. Today, AI has started to advance natural sciences by improving, accelerating, and enabling our understanding of natural phenomena at a wide range of spatial and temporal scales, giving rise to a new area of research known as AI for science (AI4Science). Being an emerging research paradigm, AI4Science is unique in that it is an enormous and highly interdisciplinary area. Thus, a unified and technical treatment of this field is needed yet challenging. This work aims to provide a technically thorough account of a subarea of AI4Science; namely, AI for quantum, atomistic, and continuum systems. These areas aim at understanding the physical world from the subatomic (wavefunctions and electron density), atomic (molecules, proteins, materials, and interactions), to macro (fluids, climate, and subsurface) scales and form an important subarea of AI4Science. A unique advantage of focusing on these areas is that they largely share a common set of challenges, thereby allowing a unified and foundational treatment. A key common challenge is how to capture physics first principles, especially symmetries, in natural systems by deep learning methods. We provide an in-depth yet intuitive account of techniques to achieve equivariance to symmetry transformations. We also discuss other common technical challenges, including explainability, out-of-distribution generalization, knowledge transfer with foundation and large language models, and uncertainty quantification. To facilitate learning and education, we provide categorized lists of resources that we found to be useful. We strive to be thorough and unified and hope this initial effort may trigger more community interests and efforts to further advance AI4Science.

1 INTRODUCTION

This work presents a unified technical review of AI for quantum, atomistic, and continuum systems, organized across physical scales and connected by shared challenges including symmetry, interpretability, OOD generalization, foundation models, and uncertainty quantification. It also provides educational resources, software and benchmarks, while remaining a selective and evolving survey.

  • 1 INTRODUCTION: Deep learning is advancing AI for science by improving, accelerating, and enabling scientific understanding across spatial and temporal scales, creating a new paradigm of interdisciplinary discovery.Deep learning can accelerate solutions to Schrödinger and Navier–Stokes equations and, once trained on simulator data, predict faster than simulators while supporting broader OOD settings.
  • 1.1 Scientific Areas: The review covers seven scientific domains—quantum mechanics, DFT, small molecules, proteins, materials, molecular interactions, and continuum mechanics—organized by the spatial and temporal scales they model.It surveys neural wavefunctions, quantum-tensor prediction, molecular and protein representation and generation, crystal-property prediction and design, molecular interactions, and PDE surrogate modeling.
  • 1.2 Technical Areas of AI: Symmetry and equivariance provide the main unifying thread across the review, alongside interpretability, OOD generalization and causality, foundation and large language models, uncertainty quantification, and education.These challenges reflect the need to encode physical invariances, understand governing rules, handle distribution shifts, learn with limited labels, and support robust decisions.
  • 1.3 Integrative Multi-Scale Analysis: The survey integrates quantum, atomistic, and continuum analyses across scales from picometers to kilometers, with theoretical-level choices determined by the phenomena and computational complexity of interest.Quantum mechanics, DFT, molecular dynamics, and PDEs address progressively larger systems and scales, and analyses at different levels can inform one another.
  • 1.4 Online Resources: The work supports continued learning and community development through an online portal, maintained mindmap, feedback channel, software library, benchmarks, and curated literature and resources.The portal and AIRS repository are intended to receive regular updates and community contributions.
  • 1.5 Scope and Feedback: The survey is selective rather than comprehensive or conclusive, focusing on AI for quantum, atomistic, and continuum systems while expecting methods and benchmarks to expand as the field develops.The authors acknowledge that relevant work may be omitted and invite community feedback through the online portal.
  • 1.7 Notations: The review establishes shared mathematical notation for scalars, vectors, matrices, particle systems, transformations, Dirac notation, and domain-specific formulations, with additional notation summarized in Table 1.Table 1 identifies notations used within individual scientific areas.

2 SYMMETRIES, EQUIVARIANCE, AND THEORY … 2.4 Equivariant Data Interactions

The section develops symmetry-aware learning from foundational motivations through discrete and continuous equivariance, geometric featurization, and equivariant data interactions. It emphasizes tensor products as necessary for strict SE(3)-equivariance while contrasting them with efficient but approximately equivariant spherical-channel methods.

  • 2.1 Overview: Symmetry-aware representations remove dependence on arbitrary coordinate choices, while architecture-level constraints address the limitations of data augmentation and support robust scientific prediction.Data augmentation requires extra capacity, cannot fully cover infinitely many transformations, and may not preserve equivariance through deep models; symmetry-adapted architectures avoid these drawbacks.
  • 2.2 Equivariance to Discrete Symmetry Transformations: Group-equivariant convolutions achieve equivariance to discrete rotations by rotating kernels, tracking permutation-structured feature responses, and pooling over the rotation axis.For 90°, 180°, and 270° rotations, lifting and group convolutions preserve equivariance across layers, while final pooling handles feature-map permutations.
  • 2.3 Equivariant Featurization of 3D Geometries: For continuous SE(3) transformations, spherical harmonics encode relative 3D geometry into features that transform through Wigner-D matrices at different rotation orders.Order ℓ=0 represents invariant properties such as energy, whereas ℓ=1 represents vector-like quantities such as force fields.
  • 2.4 Equivariant Data Interactions: Equivariant data interactions combine spherical-harmonic geometry features with node features so messages and updated representations transform consistently under SE(3).The section reviews operations designed to preserve equivariance when geometric features interact with high-dimensional features on molecular or graph structures.
  • 2.4.1 Equivariant Data Interactions via Tensor Product.: Tensor-product message passing uses distance-dependent filters, spherical harmonics, and Clebsch-Gordan matrices to construct equivariant messages over neighboring nodes.Multiple tensor-product layers can be stacked, and features with different rotation orders can be combined to form richer node representations.
  • 2.4.1 Equivariant Data Interactions via Tensor Product.: Spherical-harmonic tensor products are not only sufficient but strictly necessary for SE(3)-equivariance in the reviewed interaction framework.This necessity result is attributed to Weiler et al. [2018].
  • 2.4.2 Approximately Equivariant Data Interactions via Spherical Channel Networks.: Spherical channel networks update equivariant features nonlinearly by converting spherical-harmonic coefficients to sampled spherical functions, applying pointwise convolutions, and converting back.eSCN reduces tensor-product cost by aligning each edge with a coordinate axis, exploiting sparse Clebsch-Gordan structure, and applying an SO(2) convolution.
  • 2.4.2 Approximately Equivariant Data Interactions via Spherical Channel Networks.: SCN and eSCN are only approximately equivariant because exact preservation requires rotations matching the finitely sampled spherical grid; scalarization offers another efficiency strategy.Scalarization computes high-order interactions as scalars before restoring higher-order representations through attention or gating, balancing efficiency and expressiveness.

2.5 Intuitive Physics and Mathematical Foundations

This section develops intuitive foundations for equivariant learning by connecting group representations, tensor-product decompositions, spherical-harmonic projections, and angular-momentum concepts. Examples show how symmetry constraints structure learned features and reduce allowable parameters.

  • 2.5.1 Overview.: Changing basis decomposes polynomial features into irreducible representations, establishing that predictably transforming learned spaces must be composed of irreps.For 3D quadratic features, the decomposition separates invariant L0 and rotational L2 subspaces; tensor products combine representations and decompose them into irreps.
  • 2.5.2 Illustration of Irreducible Representations via A Discrete Example.: A square’s four-fold rotations separate features into invariant scalar A, sign-changing scalar B, and rotating vector E irreps, constraining equivariant outputs.An invariant output must be proportional to A, while a vector output must be proportional to E; symmetry can reduce a quadratic vector function to only two independent weights.
  • 2.5.3 Tensor Products and Clebsh-Gordan Coefficients.: A 3D vector tensor product forms a 9-dimensional bilinear space that decomposes under SO(3) into stable 1D, 3D, and 5D subspaces.The dot product spans the 1D scalar subspace, the cross product spans the 3D vector subspace, and Clebsch-Gordan coefficients change basis to these irreducible blocks.
  • 2.5.4 Spherical Harmonics Projections and Equivariant Networks. 𝑚:: Spherical harmonics provide complete orthogonal bases for sphere functions, while truncated projections encode vector directions and increasingly approximate spherical Dirac deltas as ℓmax grows.Equivariant descriptors combine a radial basis for magnitude with spherical-harmonic expansions for direction; Figure 7 illustrates progressively thinner reconstructions for ℓmax = 2, 4, 6, 8, and 10.
  • 2.5.5 Spherical Harmonics Functions and Angular Momentum.: Spherical harmonics and tensor products connect equivariant representations to quantum mechanics through hydrogenic wavefunctions and angular-momentum coupling.Hydrogenic electron wavefunctions factor into radial functions and spherical harmonics, while coupled-electron angular momentum is obtained from tensor products of separate angular momenta.

2.6 Group and Representation Theory

This section introduces the group- and representation-theoretic language underlying equivariant neural networks, from symmetry groups and their actions to representations and irreducible components. These concepts formalize how symmetries transform data, features, and network outputs.

  • 2.6.1 Symmetry Groups.: Groups formalize symmetry transformations through composition, inverses, identity, and associativity; commuting groups are abelian, while matrix groups and subgroups provide geometric examples.Planar rotations are abelian, whereas general 3D rotations, translations, reflections, and permutations need not commute; SO(n) contains rotations and O(n) additionally contains reflections.
  • 2.6.2 Group Actions and Equivariant Maps.: Group actions describe how a symmetry group transforms a set or space, while invariant functions produce unchanged outputs under transformed inputs.In deep learning, image classification should often be translation invariant, and a molecule’s ionization energy should be invariant under rotations and reflections.
  • 2.6.3 Group Representations.: Group representations describe symmetry actions on vector spaces as invertible matrices or linear maps that preserve group composition, inverses, and identity.For SO(3), a 3D Cartesian vector uses 3 × 3 rotation matrices, parameterized for example by axis-angle, Euler angles, or quaternions.
  • 2.6.4 Irreducible Representations.: Representations may be isomorphic under a change of basis and reducible when they contain invariant subspaces; irreducible representations form independent components that can be used as natural data types in equivariant networks.For most groups, representations can be decomposed into irreducibles, making equivariance constraints easier to solve via Schur’s lemma, although architectures such as group convolutions need not explicitly use irreps.
  • 2.6.4 Irreducible Representations.: Irreducibles can be labeled by characters for finite groups and by highest-weight data for infinite groups such as SO(3), where the label is the angular-momentum quantum number ℓ.Equivalent representations require a choice of matrix realization, while character tables or dominant integral weights provide labels less dependent on that choice.

2.7 𝑆𝑂(3) Group and Spherical Harmonics · 2.8 A General Formulation of Equivariant Networks via Steerable Kernels

The framework represents feature fields through group actions and derives all linear equivariant maps as convolutions with symmetry-constrained steerable kernels. This representation-theoretic view unifies existing equivariant architectures, supports mixed field types, and connects spherical-harmonic constructions to general SO(3)-equivariant kernels.

  • 2.8.1 Feature Vector Fields.: Feature vector fields assign feature vectors to spatial points, with their transformation laws determined by a G-representation ρ acting on individual features.Scalar and tangent vector fields arise from the trivial and defining representations, while tensor, irreducible, and regular representations describe broader feature types.
  • 2.8 A General Formulation of Equivariant Networks via Steerable Kernels: Steerable CNNs generalize equivariant networks to affine groups Aff(G), combining translations with rotations, reflections, scaling, or shearing through feature vector fields.Translations necessitate convolution operations, while G specifies additional linear transformations in Euclidean space.
  • 2.8.2 Steerable Convolutions.: Every linear equivariant map between fields of types ρ_in and ρ_out is a convolution whose matrix-valued kernel satisfies an additional G-steerability constraint.Convolution provides translational equivariance, whereas steerability enforces compatibility with the input and output field representations.
  • 2.8.2 Steerable Convolutions.: Steerable kernels form a linear subspace, so implementations can expand them in a symmetry-respecting basis with learnable coefficients.Such bases exist for SO(3) irreducible representations, groups G≤O(2), and arbitrary compact groups, with implementations available in escnn.
  • 2.7 𝑆𝑂(3) Group and Spherical Harmonics: For SO(3), kernels mapping scalar inputs to irreducible outputs have spherical-harmonic angular parts and freely learnable radial parts, explaining corresponding tensor-product operations.Mappings between two SO(3) irreducible representations are characterized through Clebsch–Gordan decomposition and recover the general tensor-product operation.
  • 2.8.2 Steerable Convolutions.: The formulation fixes group actions first and derives complete equivariant linear maps, revealing when prior architectures use only a subset of admissible kernels.This can expose unnecessary restrictions on expressive power, while relating tensor-product operations one-to-one with steerable-kernel solutions.
  • 2.8.2 Steerable Convolutions.: Steerable CNNs support hybrid models that simultaneously combine regular, irreducible, quotient, and other field types within their feature spaces.This extends beyond approaches restricted to a single field type or class of field types.
  • 2.8.2 Steerable Convolutions.: The representation-theoretic framework extends to homogeneous spaces, Riemannian manifolds, local gauge transformations, quantum-mechanical tensor operators, and steerable partial differential operators.These extensions connect steerable kernels to spherical convolutions, gauge field theory, the Wigner–Eckart theorem, and symmetry-respecting physical-science operators.

2.9 Open Research Directions

The section identifies open directions for making symmetry-aware scientific AI more expressive, scalable, and robust when exact symmetries fail or outputs break symmetry. It highlights unresolved symmetry breaking, computational costs, universality, frame-based alternatives, and approximate equivariance.

  • 2.9.1 Symmetry Breaking.: Equivariant models cannot predict a single lower-symmetry outcome from symmetric inputs, while representing or sampling all degenerate outcomes remains open.When one outcome is desired, loss gradients can identify missing symmetry-breaking input [Smidt et al. 2021]; when all outcomes are valid, existing methods do not yet represent them properly.
  • 2.9.2 Empirical Benefits and Expense of Equivariance versus Invariance.: Equivariance improves performance over invariance as feature order increases, but tensor-product contractions dominate computational expense and motivate optimized algorithms, kernels, compilers, and hardware.These costs can be reduced in voxel models through precomputed contractions and may be addressed by eSCN-like operations [Passaro and Zitnick 2023].
  • 2.9.3 Universality of Equivariant Neural Architectures.: Many equivariant architecture families are universal, but graph neural networks are generally not, making expressivity analysis an important criterion for architecture selection.Universality has been established for polynomial invariant/equivariant networks [Yarotsky 2018], tensor-product Lie-group architectures [Bogatskiy et al. 2022], and SE(3)-transformers and tensor field networks [Dym and Maron 2021], while geometric graph expressivity remains active research [Xu et al. 2018; Joshi et al. 2023].
  • 2.9.4 Frame Averaging as an Alternative for Equivariance.: Frame averaging offers model-agnostic equivariance with efficient standard architectures, but frame selection involves smoothness, computational, permutation, and discontinuity trade-offs.Tensor-field-style architectures can scale as O(L^6), whereas frame methods avoid group-specific architectural specialization; promising directions include learned or minimal frames, but PCA-based frames can break permutation symmetry and some frames are discontinuous.
  • 2.9.5 Approximate Equivariance.: Approximate-equivariant networks relax exact symmetry constraints to handle boundaries, noise, discretization, or inherent symmetry breaking, improving performance on approximate-symmetry tasks without significant degradation elsewhere.Residual Pathway Priors [Finzi et al. 2021] and relaxed group convolutions outperform perfectly equivariant and generic architectures in reported approximate-symmetry settings, but their long-term advantage under genuine symmetry remains unresolved.

3 AI FOR QUANTUM MECHANICS

This section presents neural-wavefunction methods for learning quantum ground states in spin systems, many-electron systems, molecules, and periodic fermionic systems. It covers variational and diffusion Monte Carlo, antisymmetry and symmetry, orbital and transferable architectures, optimization and sampling, energy-surface modeling, and challenges in implicit modeling, size consistency, and computational scaling.

  • 3.1 Overview: Variational Monte Carlo learns neural quantum states by sampling configurations from the wavefunction-induced distribution, estimating energy, and updating parameters to minimize it.The variational principle guarantees the estimated energy is no lower than the ground-state energy; optimized low-energy wavefunctions approximate the ground state.
  • 3.2.2 Technical Challenges.: Neural quantum states for spin systems must capture symmetries, difficult sign structures, and diverse lattice geometries, which strongly affect ground-state learning.Symmetry-aware architectures reduce the hypothesis space and improve data efficiency, while frustrated systems make sign learning especially difficult.
  • 3.2.3 Existing Methods.: Spin-system wavefunctions use restricted Boltzmann, feed-forward, convolutional, graph, autoregressive, and recurrent networks, with graph and lattice-convolution approaches extending applicability beyond regular grids.Symmetry averaging, group-equivariant convolution, complex-valued or separate amplitude-phase models, and universal multi-geometry ansätze address key structural challenges.
  • 3.2.4 Optimization Methods.: Spin-system optimization uses gradient descent, stochastic reconfiguration, and imaginary-time supervised wavefunction optimization, but direct stochastic-reconfiguration inversion is expensive for large networks.Iterative solvers reduce stochastic-reconfiguration complexity, while MinSR-style reformulations scale linearly with variational parameters and suit small batch dimensions.
  • 3.2.5 Datasets and Benchmarks.: Spin-system models are trained concurrently on dynamically sampled configurations and commonly evaluated by energy, while open problems include expressiveness, comprehensive benchmarking, and correlated MCMC samples.Autoregressive wavefunctions offer a potential alternative by bypassing MCMC and supporting more efficient exact sampling.
  • 3.3 Learning Ground States for Many-Electron Systems: Many-electron ground-state learning operates in continuous electron-coordinate space and must satisfy fermionic antisymmetry, motivating first-quantization neural wavefunctions alongside discrete antisymmetric basis methods.Continuous-space methods avoid dependence on a chosen basis and complement second-quantization approaches used for molecules.
  • 3.3.1 Problem Setup.: Many-electron ground states are learned by minimizing Monte Carlo-estimated energy subject to same-spin fermion antisymmetry, with the variational principle providing an upper bound on ground-state energy.The formulation extends to generalized wavefunctions conditioned on molecular structures, enabling joint learning of potential energy surfaces across geometries.
  • 3.3.2 Technical Challenges.: The central challenges are enforcing antisymmetry, modeling expressive interacting-electron orbitals, optimizing to chemical accuracy, and handling multiple geometries while preserving energy symmetries and size consistency.Chemical accuracy is 1 kcal/mol (1.594 mEh or 0.043 eV), and N2 energy errors must be below 0.2% for chemical applications.
  • 3.3.3 Existing Methods.: Slater determinants enforce fermion antisymmetry by changing sign when same-spin electron rows are exchanged, while sums of determinants and alternative constructions increase expressiveness or reduce cost.The surveyed alternatives include pairwise antisymmetric products, hidden fermions, permutation sums, and antisymmetric geminal power wavefunctions for superfluids.
  • 3.3.3 Existing Methods.: Neural orbital models capture collective electron interactions with permutation-equivariant features, distance-based convolutions, mean-field or attention mechanisms, while envelopes, cusp terms, Jastrow factors, and classical orbitals encode physics.Periodic systems use periodic coordinate embeddings or envelopes, and FermiNet, PauliNet, FermiNet+SchNet, PsiFormer, Moon, and MP-NQS instantiate these design choices.
  • 3.3.3 Existing Methods.: Variational Monte Carlo uses unbiased Hermitian-property gradients from MCMC samples, while natural-gradient or stochastic-reconfiguration optimization uses KFAC or conjugate-gradient approximations to make large models tractable.Metropolis-Hastings proposals are tuned to an acceptance ratio around 0.5, and second-order optimization is described as critical for accurate optimization.
  • 3.3.3 Existing Methods.: Generalized wavefunction methods adapt orbitals across geometries or compounds using weight sharing, nuclear-conditioned networks, transferable atomic orbitals, and graph-learned embeddings.PESNet avoids per-structure retraining but does not transfer across atom sets; TAO transfers better to new molecules, while Globe learns multiple compounds’ ground states simultaneously.
  • 3.3.3 Existing Methods.: Generalized wavefunctions enforce Euclidean and nuclear-permutation symmetries, size consistency, and efficient energy-surface prediction through equivariant frames or augmentation, decaying receptive fields, and temporal surface averaging.Invariant wavefunctions restrict the function class, whereas Moon uses decaying spatial filters for size consistency and PlaNet averages fitted energy surfaces over training time.
  • 3.3.4 Datasets and Benchmarks.: Quantum-wavefunction benchmarks sample atom coordinates from the neural-wavefunction distribution and assess average energy and standard deviation, with lower energy indicating greater accuracy.Test systems include atoms, molecules, special configurations, compound structures, and transition energies such as cyclobutadiene automerization and Fe ionization [Spencer et al. 2020; von Glehn et al. 2023].
  • 3.3.5 Open Research Directions.: Current many-electron methods remain limited by antisymmetry optimization, explicit real-space wavefunction modeling, and O(N^4) computational complexity, restricting calculations to at most 80 electrons.Future scaling depends on more efficient sampling, improved optimization, stronger weight sharing, and implementation accelerations such as Jax.

4 AI FOR DENSITY FUNCTIONAL THEORY

This section introduces AI methods for density functional theory, covering equivariant quantum-tensor networks, datasets, and machine-learned density functionals grounded in physical constraints. It surveys approaches to improving accuracy, generalization, convergence, and computational efficiency while outlining open challenges in physical reliability, optimization, broader validation, and data generation.

  • 4.1.3 Density Functional Theory In Practice.: DFT replaces exponentially scaling wavefunction calculations with basis-set and self-consistent-field approximations, but iterative SCF calculations remain expensive for large systems.Deep learning can directly predict converged Hamiltonian matrices, eliminating iterative optimization, while the unknown exchange-correlation functional remains the key DFT challenge.
  • 4.1.1 Density Functional Theory.: DFT maps interacting many-body systems to noninteracting Kohn-Sham electrons, with ground-state density determining the external potential and total energy minimized by the ground-state density.The Kohn-Sham energy comprises kinetic, Hartree, external-potential, and exchange-correlation terms, while the exact exchange-correlation functional remains unknown.
  • 4.1.2 The Kohn-Sham Equation.: Approximate exchange-correlation functionals cause delocalization, self-interaction, and static or dynamical correlation errors, motivating machine-learned functionals subject to physical constraints.Hybrid functionals partly reduce self-interaction error through exact exchange but cost more than typical GGA and LDA methods; advanced correlation treatments are computationally expensive.
  • 4.2 Quantum Tensor Learning: Quantum tensor learning predicts Hamiltonian matrices from atom types and coordinates because these tensors determine physical properties including total energy, charge density, and electric polarization.The Hamiltonian matrix is represented in an electronic-orbital space whose size varies with the chemical elements in the system.
  • 4.2.2 Technical Challenges.: Quantum tensor models must guarantee permutation, translation, and rotation equivariance while remaining flexible across variable tensor sizes and computationally efficient.Under rotation, a Hamiltonian block transforms as D_l1(R) B D_l2(R^T), requiring architectures that construct equivariant matrices rather than only equivariant features.
  • 4.2.3 Existing Methods.: Existing quantum tensor networks combine node-wise message passing, pairwise feature construction, and matrix assembly; SchNorb and DeepH use invariant features, whereas PhiSNet and QHNet enforce equivariance with spherical harmonics and tensor products.Tensor products increase computational cost, and QHNet improves efficiency over PhiSNet by reducing their number.
  • 4.2.3 Existing Methods.: Equivariant quantum-tensor networks construct Hamiltonian matrices from pairwise features using tensor expansion, orbital selection, basis transformations, or coordinate transformations to preserve required symmetries.QTNet and PhiSNet use tensor expansion, DeepH applies inverse Wigner transformations, and spin or time-reversal issues require basis transformations.
  • 4.2.4 Datasets and Benchmarks.: 83.12 × 10−6𝐸ℎ MAE on mixed MD17 shows QHNet trained on QH9 can predict Hamiltonian matrices across unseen molecules.QH9 includes static and dynamic datasets designed to evaluate in-distribution and out-of-distribution molecular generalization.
  • 4.2.5 Open Research Directions.: Quantum tensors can provide physics-based features for molecular-property and force-field models, but obtaining accurate tensors efficiently remains an open challenge.Low-energy systems often have atomic-orbital-like wavefunctions, enabling accurate low-cutoff representations and applications including materials discovery.
  • 4.3 Density Functional Learning.: Machine learning density-functional approaches predict exchange-correlation, kinetic, universal, correction, or other functionals using numerical or symbolic models, sometimes incorporating established functionals or exact constraints.Approaches may learn from scratch, refine existing functionals, or impose physical constraints analytically or through constrained training data.
  • 4.3.1 Machine Learning Exchange-Correlation Energy Functionals.: ML exchange-correlation functionals improve accuracy by combining data-driven learning with exact physical constraints, although data-only models may require many samples and generalization remains difficult.Constraints can be imposed analytically or through training data; examples address fractional charge and spin, asymptotics, scaling, Lieb-Oxford bounds, and derivative discontinuities.
  • 4.3.2 Machine Learning Kinetic Energy Functionals.: ML kinetic-energy functionals enable orbital-free DFT with quasi-linear scaling, but noisy functional derivatives complicate optimization; reported models achieve chemical accuracy in several molecular and material tests.The kinetic-energy approximation has a larger energetic impact than exchange-correlation errors, while constraint enforcement improves some cases but not uniformly.
  • 4.3.3 Datasets and Benchmarks.: High-level datasets such as ACCDB and SOL62 support training and benchmarking machine-learned functionals, which increasingly approach chemical accuracy at lower cost than higher-level calculations.ACCDB contains 10,049 structures and 8,656 unique reference data points, or 44,931 including reaction energies, across varied chemical tasks.
  • 4.3.4 Open Research Directions.: Symbolic regression, evolutionary search, and AFE+DQN offer routes to discover or improve analytical density functionals while incorporating mathematical libraries, physical constraints, and data-driven feature selection.Ma et al. reconstructed known functionals and developed GAS22 with improved MGCDB84 error; AFE+DQN improved classification and regression scores with less computation time on three materials databases.
  • 4.3.4 Open Research Directions.: Physically constrained training improves machine-learned density functionals, while Bayesian optimization can predict Hubbard U more accurately than conventional linear response for strongly correlated materials.DM21 and DM21mu, trained on constrained data, significantly outperform unconstrained DM21m; exact output constraints and physically constrained kinetic-energy functionals remain open directions.
  • 4.3.4 Open Research Directions.: Gradient-free optimization, richer training data beyond converged densities and energies, and correction of non-uniform sampling are key unresolved challenges for machine-learned kinetic-energy functionals.Monte Carlo optimization avoids evaluating functional derivatives but still needs extension to three dimensions; KS regularizers and automatic differentiation broaden training signals.
  • 4.3.4 Open Research Directions.: Machine-learned density functionals need broader validation across solids and chemically diverse molecules, alongside more efficient KS-DFT optimization and more accurate datasets generated by quantum-mechanical AI/ML.Proposed directions include LC20 and SOL62 datasets, stochastic-gradient optimization with embedded wave-function orthonormality, and applications spanning chemistry, materials, engineering, biology, and pharmaceuticals.

5 AI FOR SMALL MOLECULES

This section surveys AI for small molecules as a 3D-centered field, covering symmetry-aware representation learning, conformer and molecule generation, molecular dynamics, stereochemistry, and conformational flexibility. It emphasizes challenges in expressivity, chemical validity, data, efficiency, benchmarking, and reliable modeling of flexible molecular geometries.

  • 5.1 Overview: AI for small molecules centers on five 3D tasks: representation learning, conformer generation, molecule generation from scratch, molecular dynamics simulation, and learning stereochemistry and conformational flexibility.Figure 17 organizes representative invariant, equivariant, generative, dynamics, stereochemical, and conformer-ensemble methods across these tasks.
  • 5.2 Molecular Representation Learning: Molecular representation learning uses 3D point clouds to support molecule-level and atom-level predictions, with learned backbones enabling applications such as drug discovery and material design.A molecule is represented by atom types and coordinates, while scalar or tensorial targets must transform as invariant or equivariant quantities under reference-frame changes.
  • 5.2.2 Technical Challenges.: 3D molecular representations must match target symmetries, distinguish enantiomers and conformers, and remain efficient enough for scalable training and inference.Energy prediction requires SE(3)-invariant representations, whereas per-atom force prediction requires SO(3)-equivariant representations.
  • 5.2.3 Overview of Existing Methods.: Existing 3D molecular GNNs are organized by tensor order and body order, ranging from invariant scalar features to order-1 vectors, higher-order tensors, and many-body interactions.Figure 19 identifies tensor order and body order as key design choices, while equivariant architectures require operations such as tensor products to preserve SE(3) equivariance.
  • 5.2.4 Invariant Methods (ℓ= 0 Scalar Features).: Invariant models trade geometric discriminative power against efficiency: SchNet has O(nk) complexity, whereas DimeNet and GemNet increase complexity to O(nk^2) and O(nk^3) while incorporating higher-order geometry.SphereNet retains O(nk^2) complexity through reference-node selection, while ComENet claims complete 3D distinguishability with O(nk) complexity.
  • 5.2.5 Equivariant Methods (ℓ= 1 Vector Features).: Equivariant models preserve geometric transformations through restricted vector operations or higher-order tensor products, with attention and specialized nonlinearities expanding their expressive architectures.Order-1 methods use operations such as scaling, summation, linear transformation, scalar products, and cross products; higher-order methods use spherical tensors, tensor-product messages, equivariant linear layers, and radial-function–spherical-harmonic filters.
  • 5.2.6 Equivariant Methods (ℓ≥1 Tensor Features).: Equivariant models aggregate neighbor messages with sum operations or attention, while update designs differ in residual connections that help retain chemical information such as atom type.SE(3)-Transformer uses dot-product attention, whereas Equiformer uses MLP attention computed only from ℓ=0 features to preserve equivariance.
  • 5.2.7 Higher Body Order Methods.: Higher-body-order methods use atomic cluster expansion to capture many-body interactions with computational cost linear in the number of neighbors, enabling models such as Linear ACE, MACE, Allegro, and PACE.ACE replaces expensive neighbor-combination sums with products of atomic basis functions; these methods differ in layering, edge-centered operations, and convolution construction.
  • 5.2.8 Model Outputs.: Molecular outputs commonly predict energy and derive forces by differentiating energy with respect to coordinates, preserving energy conservation; equivariant methods are more direct for targets such as Hamiltonian matrices.Higher-body-order models first compute per-atom or pairwise local energies and sum them to obtain molecular energy.
  • 5.2.9 Datasets and Benchmarks.: Common benchmarks include QM9, MD17 and its revised or CCSD(T) variants, ISO17, and Molecule3D, evaluated primarily with mean absolute or mean square error on molecular properties, energies, and forces.QM9 contains more than 130k molecules, MD17-family datasets provide conformations with energies and forces, ISO17 spans 129 isomers, and Molecule3D contains around 4 million molecules.
  • 5.2.10 Open Research Directions.: Open directions include jointly learning from 2D and 3D information, pre-training, improving provable geometric expressivity, and making higher-order equivariant models efficient enough for larger biomolecules and datasets.Expressivity depends on depth, tensor order, and scalarization body order, but higher-order computation and geometric oversquashing limit practical networks; invariant models offer a scalability trade-off.
  • Recommended Prerequisites: Section 5.2: Three-dimensional molecular geometries improve property prediction and support applications such as molecular dynamics and docking, but their DFT-based acquisition is computationally expensive, motivating machine-learning reconstruction.Spatial configuration can substantially affect molecular properties, including the behavior of isomers and antibody binding sites.
  • 5.3.1 Problem Setup.: Molecular geometry reconstruction separates 3D geometry generation, which models p(C|G) for low-energy conformers, from 3D geometry prediction, which estimates the equilibrium ground-state geometry from a 2D graph.A 2D molecule is represented by atom types and edge types, while the 3D molecule additionally contains atom coordinates; lower-energy geometries are generally more stable.
  • 5.3.2 Technical Challenges.: Conformer reconstruction must ensure geometrical validity, chemical validity, and SE(3)-invariant position distributions so that symmetry-related atoms remain distinct and rotated or translated conformers are equivalent.Challenges include valid Euclidean distance matrices, planar aromatic rings or π bonds, non-planar rings and macrocycles, and the condition p(RC + t1^T|G) = p(C|G).
  • 5.3.3 Existing Methods.: Current 3D molecule-generation methods use score matching or diffusion with E(3)-equivariant/invariant modules, torsional diffusion that preserves local structures, or predictive models for geometrically valid conformers.Table 19 distinguishes coordinate, distance, torsion, and predictive strategies; torsional diffusion retains RDKit-generated bond lengths and angles but cannot refine certain ring structures.
  • 5.3.4 Datasets and Benchmarks.: Conformer benchmarks span small, rigid GEOM-QM9 molecules and larger, flexible GEOM-Drugs molecules, whose mean and maximum rotatable bonds reach 6.5 and 53, respectively.GEOM-QM9 contains 133,258 molecules with up to 9 heavy atoms, whereas GEOM-Drugs averages 44.4 atoms and reaches 181 atoms.
  • 5.3.5 Open Research Directions.: Future conformer generation should model solvent effects, reduce data requirements through transfer or few-shot learning, and generate high-energy transition-state structures rather than only low-energy conformers.These directions address vacuum-only benchmarks, limited data for specialized compounds, and the importance of transition states for reaction kinetics and mechanisms [Choi 2023; Duan et al. 2023a].
  • 5.4.2 Technical Challenges.: Generating molecules from scratch requires SE(3)-invariant distributions over atom types and coordinates, either by directly modeling coordinates or by generating complete invariant distances, angles, and torsions.Direct coordinate methods zero-center molecules and use CoM-free Gaussian distributions, while feature-based methods must ensure valid, complete structures that permit coordinate reconstruction.
  • 5.4.3 Existing Methods.: Representative from-scratch methods include E-NFs, EDM, GeoLDM, EDMNet, G-SchNet, and G-SphereNet; only EDM supports molecular-property conditioning, while others use implicit property optimization.E-NFs, EDM, and GeoLDM directly generate coordinates, whereas EDMNet and autoregressive methods generate invariant geometric features; Table 20 summarizes these distinctions.
  • Recommended Prerequisites: Section 5.2: ML molecular dynamics replaces expensive quantum energy-and-force calculations with learned force fields, while enhanced sampling and coarse-graining target longer timescales and reduced complexity.MD trajectories yield observables such as radial distribution functions, mean-squared displacement, diffusivity, and free-energy surfaces; exact trajectory recovery is unnecessary because MD is chaotic.
  • 5.5.2 Technical Challenges.: Force-field accuracy does not guarantee reliable simulations: nonsmooth potential-energy surfaces, rare events, and femtosecond integration steps make stability, long-timescale sampling, and observable prediction difficult.A benchmark found misalignment between force/energy prediction and simulation performance, showing that force and energy error alone are insufficient evaluation protocols.
  • 5.5.3 Existing Methods.: ML force-field research combines symmetry-aware architectures with reaction-coordinate discovery, learned coarse-grained dynamics and mappings, and rare-event sampling, but must improve long-range interactions, efficiency, stability, and information retention.Key benchmark systems include MD17, alanine dipeptide, Chignolin, water, electrolytes, and polymers; future methods emphasize active learning, parallelizable models, multiscale coarse-graining, and transition-path or generative sampling.
  • Recommended Prerequisites: Section 5.2: Stereoisomers share 2D graphs but can differ chemically, while conformer ensembles and thermodynamic averages determine many observable molecular properties.Stereochemistry includes tetrahedral, axial, helical, and E/Z forms; conformers interconvert at environment- and temperature-dependent rates, and properties may depend on ensembles or higher-energy active geometries.
  • 5.6.1 Problem Setup.: Conformer-ensemble learning represents discrete geometries with energies, subsets, Boltzmann weights, and tasks including stereoisomer classification, averaged-property prediction, and active-conformer identification.The ensemble is discretized using an RMSD threshold, with subsets reflecting relevant interconversion timescales and weights determined by expected experimental presence.
  • 5.6.2 Technical Challenges.: Modeling stereochemistry and flexibility is difficult because 2D features miss geometry, simple 3D invariant networks may miss chirality, and multiple conformers increase optimization and computational costs.High-quality conformers can require expensive simulations at inference time; cheaper conformers may reduce accuracy, while random conformers may omit property-active structures and realistic solvent effects.
  • 5.6.3 Existing Methods.: Existing methods encode chirality with graph or SMILES features, permutation-aware message passing, torsion-based ChIRo, or 3D GNNs, while conformer ensembles commonly use multiple-instance pooling or attention.ChIRo is conformationally invariant and can use cheap RDKit conformers, but current 3D GNNs underperform 2D chiral-tag models on small-molecule chiral-property tasks; ensemble aggregation can identify important conformers but increases cost.
  • 5.6.4 Datasets and Benchmarks.: Benchmarking lacks broad stereochemistry-sensitive measurements, and existing conformer-ensemble studies show no advantage over simpler baselines despite substantially greater resources.On GEOM-DRUGS, multi-conformer models did not outperform single-conformation baselines; on a synthetic key-conformer task, random forests with ECFP4 fingerprints outperformed MIL models.
  • 5.6.5 Open Research Directions.: Future work should develop stereochemistry-sensitive datasets and models requiring fewer paired stereoisomers, extend beyond tetrahedral chirality, and efficiently transfer across conformer-generation methods.Conformer-ensemble models currently cannot outperform single-conformer models, while generating and encoding realistic ensembles remains a practical barrier for inference and virtual screening.

6 AI FOR PROTEIN SCIENCE

This section surveys AI methods and benchmarks for protein science, focusing on individual protein molecules, protein folding and representation learning, and diffusion-based protein backbone generation. It examines geometric structure and SE(3)/E(3) symmetries alongside datasets, computational and sampling challenges, and open directions.

  • 6.1 Overview: The section organizes AI for protein science around protein folding, representation learning, and backbone generation, while distinguishing coordinate, frame, and internal-angle structure representations.Protein folding methods are grouped into two-stage and end-to-end approaches; representation learning methods into invariant and equivariant networks; and backbone generation methods by their structural representation.
  • 6.2 Protein Folding: Protein folding predicts a protein’s full 3D atomic structure from its amino acid sequence, but proteins’ large size makes native-structure estimation substantially more challenging than for small molecules.The task includes backbone and side-chain coordinates, and protein molecules typically contain 1,000 to 10,000 atoms.
  • 6.2.2 Technical Challenges.: Protein-folding models must preserve SE(3) consistency, enforce physical constraints such as bond lengths and clash avoidance, and efficiently search an enormous structural space.Current methods still degrade for multi-chain proteins, proteins dissimilar to their training data or MSA databases, and complexes involving proteins, nucleic acids, or small molecules.
  • 6.2.3 Existing Methods.: AlphaFold2 established near-experimental protein-folding accuracy in CASP14, while subsequent systems target comparable accuracy with greater speed, memory efficiency, or multi-chain coverage.AlphaFold2 achieved nearly 90 GDT in CASP14; ESMFold reports up to 60× speedup by replacing MSA with a pretrained sequence language model, and OpenFold provides a memory-efficient implementation.
  • 6.3 Protein Representation Learning: Protein representation learning encodes sequence and structure for protein-level and node-level prediction, but must handle large multi-level molecules while preserving SE(3)-invariance and distinguishing non-equivalent structures.Methods reduce computational cost through atom-level hierarchical pooling or amino-acid-level graphs, while incorporating backbone, side-chain, full-atom, or surface information.
  • 6.3.3 Existing Methods.: Existing protein representations mainly use scalar or first-order directional features; higher-order and many-body methods remain an open opportunity because protein structures are large.Evaluation spans amino-acid type and protein-function prediction using datasets including CATH 4.2, Fold, Enzyme Reaction, EC, and Gene Ontology.
  • 6.3.4 Datasets and Benchmarks.: Protein-science benchmarks cover inverse folding, fold and enzyme classification, Gene Ontology prediction, and diverse biomolecular representation-learning tasks.CATH 4.2 inverse folding uses perplexity and recovery, while fold classification uses accuracy; enzyme and Gene Ontology tasks use accuracy or Fmax/AUPRpair.
  • 6.3.5 Open Research Directions.: Protein representation learning still needs better modeling of local secondary structures, tertiary spatial arrangement, surfaces, domains, and higher-order equivariance.MaSIF and dMaSIF model protein surfaces, whereas existing equivariant approaches largely remain limited to ℓ=0 and ℓ=1.
  • 6.4 Protein Backbone Structure Generation: Protein backbone generation learns pθ to approximate the real backbone distribution pPbb, enabling sampling of novel structures from the learned model.The section focuses on diffusion models, while flow-matching approaches are also noted.
  • 6.4.2 Technical Challenges.: Backbone generation must address unknown and sparse structure distributions, symmetry invariance, equivariant networks, and the vast search space of protein geometries.SE(3) invariance preserves rotations and translations without incorrectly identifying chiral proteins under reflections; zero-centering handles translations, while equivariant mappings address rotations.
  • 6.4.3 Existing Methods.: Diffusion methods trade structural granularity and computational cost across coordinates, latent spaces, angles, and residue frames, while enforcing corresponding geometric symmetries.Backbone-only generation reduces complexity; some methods subsequently predict full atom-level structures.
  • 6.4.3 Existing Methods.: Coordinate methods use zero-mean distributions and equivariant networks, Chroma models correlated constraints, FoldingDiff diffuses angles, and RFdiffusion or FrameDiff diffuse residue frames.RFdiffusion uses Brownian motion on SO(3) and self-conditioning, while FrameDiff uses manifold diffusion for rotation matrices.
  • 6.4.4 Datasets and Benchmarks.: Protein backbone generation lacks standard benchmarks; evaluations use heterogeneous structural datasets and commonly report scTM, with higher scores preferred, plus novelty.Training collections range from thousands of PDB or curated-library structures to approximately 100k examples for LatentDiff.
  • 6.4.5 Open Research Directions.: Future protein generators should produce full atom-level structures in one step and satisfy desirable properties beyond random or substructure-conditioned generation.Protein dynamics, circuit design, and mutation-effect prediction are additional open directions, supported by temporal GNNs and ProteinGym.

7 AI FOR MATERIALS SCIENCE

This section surveys AI and machine-learning methods for materials science, covering crystalline, amorphous, local, disordered, and scattering-based materials through representation learning, characterization, structure prediction, modeling, and generative design. It emphasizes symmetry-aware and equivariant approaches while addressing data scarcity, computational cost, benchmark and representation gaps, and challenges in data generation, inverse design, and scalable computation.

  • 7 AI FOR MATERIALS SCIENCE: AI for materials science spans representation learning, material generation, materials characterization, phonon calculations, and amorphous-material modeling, unified by the need to capture crystalline symmetries.Crystalline materials require permutation, E(3), and periodic transformations to be handled so equivalent representations receive identical predictions or probabilities.
  • 7.2.2 Technical Challenges.: Crystal representation learning must support unit-cell E(3) and periodic invariance while efficiently modeling large unit cells and potentially essential higher-body interactions.Current property-prediction methods mainly target regression and classification, while geometric completeness and higher-rotation-order properties remain open directions.
  • 7.2 Material Representation Learning: Crystal property prediction constructs crystal graphs and applies message-passing networks, with methods differing in whether they encode periodic information and higher-order interactions.Multi-edge graphs provide periodic and E(3) invariance from pairwise distances; Matformer, PotNet, and Ewald-MP explicitly model periodic patterns, while ALIGNN, M3GNet, and CHGNet incorporate angular or three-body information.
  • 7.3 Material Generation: Crystalline material generation must make equivalent structures equally probable under permutation, E(3), and periodic transformations, which prevents directly reusing ordinary 3D molecular-generation targets.Voxel-grid methods capture permutation and periodic invariance but not E(3) symmetries; graph-based CDVAE and SyMat capture more symmetries, with SyMat additionally achieving translation invariance by score matching pairwise distances.
  • 7.3.4 Datasets and Benchmarks: Material-generation benchmarks include Perov-5, Carbon-24, and MP-20, while practical deployment remains limited by score-matching cost and the absence of standard synthesizability metrics.These datasets contain 18,928, 10,153, and 45,231 materials, respectively, with unit-cell sizes constrained by each dataset’s construction.
  • 7.4 Materials Characterization: Materials characterization uses machine learning to predict crystal structures from experimental spectra and to reconstruct diffraction spectra from known structures, potentially accelerating workflows overwhelmed by high-volume measurements.The reviewed applications include regression of lattice parameters or atomic positions and classification of space groups, Bravais lattices, or crystal systems; the section notes that it only scratches the surface of possible methods.
  • 7.4.2 Technical Challenges.: Scattering-based crystal analysis is hindered by the phase problem, material-specific experimental variation, and the need for models that preserve crystal symmetries and generalize across simulated and experimental data.Only scattering intensities are measured, so diffraction patterns may not uniquely determine crystal structures; lower-symmetry systems and experimental data remain especially difficult.
  • 7.4.3 Existing Methods.: Existing scattering models classify crystal symmetries or regress lattice parameters from spectra, but CNN-based methods degrade for lower-symmetry systems and 1D spectral assumptions may not match the data’s symmetries.Equivariant neural networks offer a promising route for predicting one- and two-dimensional scattering spectra from crystal structures, including compressed latent-space representations of large 2D spectra.
  • 7.4.4 Datasets and Benchmarks.: Simulated scattering data are widely used because experimental benchmarks are unavailable, but simulation-to-experiment transfer is unsatisfactory; augmentation and standardized experimental datasets are needed.Relevant augmentation factors include peak shifts, broadening, texture, and noisy backgrounds.
  • 7.4.5 Open Research Directions.: Future scattering research should combine symmetry-aware representations with Fourier-space and real-space modeling, while improving performance for lower-symmetry structures and experimental data.Fourier-domain graph convolutions could be synthesized with equivariant networks to address the phase problem and connect reciprocal-space measurements with real-space structures.
  • 7.5 Local Structure and Disordered Materials Characterization: Local-structure learning targets 3D atomic configurations and tensorial chemical properties from PDF and NMR measurements, but disordered and short-range structures are difficult to characterize, simulate, and represent efficiently.Disorder can strongly affect crystalline-material properties, while amorphous systems make large DFT training datasets especially difficult to construct.
  • 7.5.3 Existing Methods.: DeepStruc successfully predicts some simple monometallic nanoparticle structures from experimental PDF data, while equivariant models outperform symmetry-invariant models by 53 % for Si chemical-shift tensors in silicates.DeepStruc uses a conditional variational autoencoder and solves an unassigned distance-geometry problem; the tensor model uses MatTEN with tensor field networks and e3nn.
  • 7.5.4 Datasets and Benchmarks / 7.5.5 Open Research Directions.: Local-structure datasets remain specialized and limited, including 3,742 simulated nanoparticle structures and 421 silicate structures with 1,387 unique silicon sites, motivating broader benchmarks.Open directions include invariant local descriptors based on spherical-harmonic spectra and ML-assisted acceleration of characterization tasks such as Bragg-peak detection and artifact identification.
  • 7.6.2 Technical Challenges / 7.6.3 Existing Methods.: Phonon prediction must address expensive and sometimes inaccurate DFT or scattering calculations, periodic and symmetric 3D inputs, and output dimensions that vary with phonon-band structure.ML interatomic potentials and equivariant geometric GNNs provide efficient alternatives: active-learning MTPs can approach DFT accuracy with fewer calculations, while VGNNs and E(3)NNs predict variable-dimensional phonon properties from crystal structures.
  • 7.7.1 Technical Approaches and Challenges.: Special quasirandom structures and cluster expansions approximate disordered configurations and support amorphous-material simulations through ATAT and CASM [Zunger et al. 1990; Wei1990SQS et al. 1990; van de Walle et al. 2002; Puchala et al. 2023].SQS constructs small configuration sets that statistically approximate random atom or defect distributions, while fitted effective Hamiltonians enable subsequent simulations.
  • 7.7.2 Existing Methods.: Machine-learning force fields support molecular-dynamics studies of amorphous materials, including Gaussian approximation potentials for amorphous carbon and silicon [Deringer and Csányi 2017; Deringer et al. 2018].MLIPs are important because molecular dynamics is a prominent tool for modeling amorphous materials.
  • 7.7.2 Existing Methods.: GNNs predict long-time glassy dynamics and propensity maps, enable inverse design of glass landscapes, and reveal geometrically ultrastable but energetically metastable states [Bapst et al. 2020; Wang and Zhang 2021].Other approaches identify hidden structural heterogeneities using autoencoders and Gaussian mixture models [Boattini et al. 2020], while generative models extrapolate structures from correlation-length kernels to arbitrary sizes [Kilgour et al. 2020; Madanchi et al. 2024b; Zhou et al. 2023b].
  • 7.7.3 Open Directions.: Amorphous-material research needs advanced inverse-design methods, especially extensions of crystalline-material diffusion and language models to large-supercell amorphous systems [Xie et al. 2022b; Jiao et al. 2023; Antunes et al. 2024; Yan et al. 2024; Zeni et al. 2025].Further priorities include efficient sampling and DFT data generation, plus computation- and memory-efficient GNNs for large-scale, long-time dynamics across multiple spatial and temporal scales.

8 AI FOR MOLECULAR INTERACTIONS

This section surveys AI methods for molecular interactions with proteins and materials, covering predictive and generative tasks, datasets, metrics, methods, and benchmarks. It emphasizes symmetry-aware structure-based design and challenges including search spaces, evaluation, structure relaxation, energy minima, periodicity, and coupled geometries.

  • 8.1 Overview: AI molecular-interaction tasks cover protein-ligand binding prediction, structure-based drug design, and molecule-material energy, force, and relaxed-structure prediction.Protein-ligand tasks include binding pose and affinity prediction plus ligand generation; molecule-material generative tasks remain unexplored.
  • 8.2.1 Problem Setup.: Docking predicts likely ligand poses, whereas binding-strength prediction estimates affinity, rankings, or binding classification under varying holo, apo, or computationally generated protein inputs.Docking may assume a known pocket or require blind search, and realistic evaluation should compare methods across protein-structure availability conditions.
  • 8.2.2 Technical Challenges.: Docking must model a high-dimensional pose distribution from sparse modes, with limited and noisy data of roughly 20k quality structural samples constraining both docking and affinity prediction.Docking is SE(3)-equivariant, while binding-strength prediction is SE(3)-invariant; protein-specific affinity models are currently useful when sufficient data exist.
  • 8.2.3 Existing Methods.: Protein-ligand docking methods progress from search-based scoring to regression and generative approaches, with DiffDock and NeuralPLexer sampling multiple poses and supporting unbound or predicted proteins.Regression methods are faster but can produce unphysical single poses, whereas generative methods model translations, rotations, torsions, contact maps, or ligand coordinates.
  • 8.2.4 Datasets and Benchmarks.: Docking evaluation uses RMSD-based pose accuracy and steric clashes, while affinity evaluation uses classification accuracy, ranking correlation, or MAE; realistic tests should use apo or predicted structures.PDBBind contains 20k complexes with 4k unique proteins, while ChEMBL provides 20 million affinity measurements without corresponding structures.
  • 8.2.5 Open Research Directions.: Docking remains unsolved, reaching 22% accuracy on ESMFold structures, while future directions include better generative structural models, joint structure-affinity learning, protein conformational changes, and statistical-mechanics-based interaction estimates.These directions aim to address noisy affinity measurements and the practical reality that binding commonly changes protein conformation.
  • 8.3 Structure-Based Drug Design: Structure-based drug design learns p(M|P) to generate 3D ligands that bind a target protein, requiring SE(3)-equivariance across the interacting protein and molecule.The task faces a chemical space estimated to exceed 10^60 possibilities, plus conformational space and drug-likeness constraints.
  • 8.3.3 Existing Methods.: Recent SBDD methods use equivariant autoregressive generation, diffusion models, virtual dynamics, or reinforcement learning to construct atoms, fragments, coordinates, or optimized molecular actions.Autoregressive methods model relative positions, while diffusion methods generate all atom coordinates and types jointly; reinforcement learning selects discrete growth or genetic-algorithm actions.
  • 8.3.4 Datasets and Benchmarks.: Structure-based drug design uses datasets including CrossDocked2020, PDBBind, DUD-E, and scPDB, with evaluation spanning molecule quality, binding affinity, and drug-likeness.CrossDocked2020 contains 22,584,102 docked complexes, while PDBBind, DUD-E, and scPDB contain 19,445, 22,886, and 16,034 protein–ligand pairs, respectively.
  • 8.3.5 Open Research Directions.: Drug-design models still struggle to generate valid, synthesizable molecules while optimizing multiple properties, motivating fragment, scaffold, and shape-based generation.DiffLinker uses an E(3)-equivariant 3D-conditional diffusion model to connect molecular fragments while conditioning on the surrounding protein pocket.
  • 8.4.1 Problem Setup.: Molecule–material interaction tasks predict energy, per-atom forces, relaxed energy, or relaxed structure from paired atomic types and coordinates.Adsorption energy is defined from adsorbate–surface, clean-surface, and gas-phase or reference-state energies.
  • 8.4.2 Technical Challenges.: These tasks require rotational invariance for energies and rotational equivariance for forces or structures, while relaxed-state prediction must model relaxation from only approximate initial structures.Energy and force predictions can be linked through F_i = −∂e/∂c_i, but direct force prediction avoids additional differentiation cost.
  • 8.4.3 Existing Methods.: Existing methods trade efficiency, expressivity, symmetry, and accuracy through direct or indirect force and relaxation pipelines, with EquiformerV2 leading indirect methods and Equiformer leading direct methods.Indirect relaxation iteratively updates positions using predicted forces and is more accurate but slower, whereas direct methods predict changes from initial to relaxed structures and are faster but less accurate.
  • 8.4.4 Datasets and Benchmarks.: OC20 and related benchmarks evaluate molecule–material prediction across in-domain and out-of-domain adsorbate or catalyst splits, while OC20-Dense targets global-minimum energies and OC22 uses total-energy targets.OC20 includes 133,934,018 S2EF training structures and 460,328 IS2RE/IS2RS training structures; test labels require Open Catalyst Project leaderboard submissions.
  • 8.4.5 Open Research Directions.: Open challenges include initial-configuration dependence, absent explicit periodicity, nonconservative direct forces, incomplete physical effects, and generation of periodic materials conditioned on molecules.Conditioned generation must model adsorption-related properties and the induced changes in molecular geometry caused by the generated material.

9 AI FOR PARTIAL DIFFERENTIAL EQUATIONS

This section surveys AI methods for partial differential equations, covering neural forward solvers for multi-resolution dynamics, long-term stability, irregular meshes, symmetries, and physical constraints, alongside inverse problems and AI-assisted inverse design. It also discusses evaluation, applications, benchmarks, and challenges including data efficiency, out-of-distribution generalization, three-dimensional scaling, and stiffness.

  • 9.1 Overview: PDE modeling describes physical systems through constraints on an unknown function and its partial derivatives, often requiring numerical solvers when closed-form solutions are intractable.Applications span airflow, weather, optical design, carbon storage, seismic propagation, materials, and volcanic activity [Pfaff et al. 2021; Li et al. 2022b; Bonnet et al. 2022].
  • 9.2 Forward Modeling: Neural PDE solvers learn mappings from initial or historical states to future states, while inverse modeling infers PDE configurations or optimizes system designs from observed solutions.Neural solvers can accelerate simulation through larger time steps, coarser spatial discretizations, direct prediction, GPU parallelism, and generalization across conditions, parameters, and geometries [Kochkov et al. 2021b; Li et al. 2021b; Gupta and Brandstetter 2023].
  • 9.2.1 Problem Setup.: Forward solvers should encode PDE symmetries because equivariant architectures and data augmentation improve generalization and sample complexity.The desired constraint is that transforming the input and then predicting agrees with transforming the prediction, expressed through equivariance to the symmetry group G.
  • 9.2.2 Technical Challenges.: Neural forward modeling must address multi-scale dynamics, multi-resolution dynamics, long-term rollout stability, symmetry preservation, and incorporation of physical principles.These challenges arise from interacting spatial scales, localized fast dynamics, accumulated prediction error, intrinsic symmetries linked to conservation laws, and statistically plausible but physically incorrect predictions.
  • 9.2.3 Existing Methods: Multi-Scale Dynamics.: Multi-scale architectures aggregate local and global information sequentially or in parallel, with neural operators such as FNO enabling resolution generalization through frequency-domain parameterization.FNO processes low- and high-frequency global and local information in parallel, while SineNet reduces latent misalignment in U-Net skip connections and improves performance under a fixed parameter budget [Zhang et al. 2024].
  • 9.2.4 Existing Methods: Multi-Resolution Dynamics.: Multi-resolution methods use non-uniform, adaptive meshes to assign fine discretization to high-gradient regions and coarse discretization elsewhere, balancing accuracy against computational cost.This is important for shockwaves, flow obstructions, and other localized dynamics whose high-gradient regions can shift during time evolution.
  • 9.2.4 Existing Methods: Multi-Resolution Dynamics.: Neural operators and graph-based models address irregular or varying discretizations through adaptive meshes, mesh-independent operators, multiscale kernels, and geometry-aware Fourier processing.MeshGraphNets and LAMP adapt meshes, while DeepONet, GKN, MGKN, Geo-FNO, and GINO target operator learning across discretizations and long-range interactions.
  • 9.2.5 Existing Methods: Long-Term Stability.: Long-term autoregressive prediction accumulates input errors because one-step models are trained on ground-truth states, motivating dissipativity constraints, noise injection, multi-step objectives, latent-space recurrence, and fewer forward passes.Stable long-horizon behavior may be evaluated through system statistics such as Fourier spectra, principal components, Pearson correlation, and rates of change rather than mean squared error alone.
  • 9.2.5 Existing Methods: Long-Term Stability.: PDE-Refiner improves rollout stability, sample complexity, and instability detection using only three sequential noise-refinement steps, while targeting low-amplitude Fourier components neglected by one-step MSE training.Its noise levels decrease according to Fourier-spectrum amplitudes during both training and inference, refining predictions in a diffusion-like procedure.
  • 9.2.5 Existing Methods: Long-Term Stability.: ACDM preserves correct fluid-flow statistics over longer rollouts and generates diverse physically consistent PDE-solution samples by denoising both the next state and previously predicted conditioning states.This conditional diffusion formulation is also relevant to uncertainty quantification because it produces diverse posterior samples.
  • 9.2.6 Existing Methods: Preserving Symmetries.: Symmetry-aware PDE models improve generalization and sample complexity by encoding exact, learned, or approximate equivariance through data augmentation, symmetry losses, equivariant convolutions, and Clifford-algebra operations.RGroup, RSteer, R-Equiv, GCAN, and CGENN accommodate symmetry breaking or geometric transformations, while CGENN reports gains across symmetry-related tasks including 3D n-body simulation.
  • 9.2.6 Existing Methods: Preserving Symmetries.: Frequency-domain and spherical equivariant architectures address irregular resolutions and global weather fields without planar projection distortions, with SFNO generating stable year-long forecasts exceeding 1,000 time steps.G-FNO supports equivariant convolutions in the frequency domain, while SWSCNNs additionally model vector fields on the sphere and scale spherical forecasting training.
  • 9.2.6 Existing Methods: Preserving Symmetries.: IsoGCN uses E(n)-equivariant message passing and an IsoAM to model differential operators efficiently, simulating heat evolution on 3D CAD meshes with more than 1 million collocation points.Its construction connects convolution, contraction, and tensor-product operations to operators including gradients, divergence, and Jacobians.
  • 9.2.7 Existing Methods: Incorporating Physics.: Physics-informed architectures encode governing laws to improve generalization and interpretability, including Hamiltonian, Lagrangian, hybrid residual, solver-acceleration, and PINN approaches.HNNs conserve learned energy in simple Hamiltonian systems, APHYNITY regularizes residual physics components, and B-PINNs provide uncertainty quantification and more accurate predictions on noisy data than PINNs.
  • 9.2.8 Datasets and Benchmarks: PDE benchmarks now span eight PDEs, one to three spatial dimensions, turbulent and fast-moving dynamics, large time steps, conditional prediction, Lagrangian systems, and irregular geometries.PDEBench includes compressible Navier–Stokes data with initial Mach numbers up to 1, while other datasets cover fluids, structural mechanics, and steady-state or time-dependent problems.
  • 9.2.9 Open Research Directions.: Current neural PDE solvers remain limited by costly training data, insufficient principled OOD-dynamics generalization, memory and optimization barriers in three dimensions, and difficulty modeling stiff nonsmooth interactions.Future work should improve sample efficiency and adaptation to unseen dynamics, scale reliably to 3D turbulent flows, and develop more general methods for stiff systems.
  • 9.3 Inverse Problem and Inverse Design: The discussion next shifts from forward PDE simulation to inverse problems that infer unknown system parameters or states from partial observations and to AI-assisted inverse design.These directions are introduced as the reverse direction of neural PDE solver modeling.
  • 9.3 Inverse Problem and Inverse Design: Inverse problems recover unknown initial conditions, boundary conditions, or PDE coefficients from observed data, whereas inverse design optimizes configurations to satisfy a predefined objective, even without an exact solution.Both can be formulated as constrained optimization, using classical or differentiable learned forward models; inverse problems assume physically plausible observations, while inverse design targets desired behavior.
  • 9.3.1 Problem Setup.: Applications span fluid grounding and assimilation, system identification, geophysical inversion, tomography, aircraft and battery design, fusion control, nanophotonics, and climate-related optimization.Examples include recovering fluid fields from sparse observations, inferring underground properties from seismic waves, and designing shapes, materials, or operating protocols.
  • 9.3.2 Technical Challenges.: Inverse problems and inverse design are limited by expensive forward modeling, adversarial nonphysical solutions, objective mismatch, ill-posedness, indirect observation, and the need to incorporate physical laws.Jointly learned forward models may require PDE or multi-objective constraints to preserve physical consistency while maintaining training efficiency and accuracy.
  • Challenges for Inverse Design: Inverse design additionally must represent hierarchical, heterogeneous design spaces, balance contradictory objectives, and adapt to objectives whose importance changes across operating conditions.Rocket composition and battery lifespan versus weight illustrate the structural and multi-objective complexity of real-world engineering design.
  • 9.3.3 Existing Methods.: Existing methods include NeRF-based reconstruction, learned graph-neural-network forward models, PINNs and hard-constrained PINNs, governing-equation discovery, adjoint optimization, differentiable simulation, and compositional diffusion models.CinDM [Wu et al. 2024] discovers formation flying for drag reduction despite training on single-airfoil dynamics, while current methods remain far from real-world engineering in physics, design, and cross-domain complexity.
  • 9.3.4 Datasets and Benchmarks.: Inverse-problem studies use simulated multi-view scenes, classical-solver rollouts, and 12 full waveform inversion datasets, whereas inverse design lacks a standard benchmark and sufficiently complex real-world tasks.The section calls for diverse benchmarks spanning more challenging physics and design spaces.
  • 9.3.5 Open Research Directions.: Future inverse-problem research should prioritize uncertainty quantification and improved training or regularization, while inverse design needs better representations, optimization for mixed discrete-continuous spaces, and broader cross-domain generality.These directions address instability in ill-posed recovery and the hierarchical, heterogeneous, diverse nature of real-world design tasks.

10 RELATED TECHNICAL AREAS OF AI

This section examines shared technical challenges in AI for science, including interpretability, out-of-distribution generalization, self-supervised and foundation models, knowledge transfer, and uncertainty quantification. It surveys approaches across molecular, protein, and continuum systems while highlighting data scarcity, domain adaptation, evaluation, scalability, and reliable scientific decision-making.

  • 10.1 Interpretability: Interpretability in AI for science focuses on instance-level explanations that identify input-graph components and patterns driving each prediction.The goal is to expose the factors underlying model outputs rather than treat geometric deep learning models as black boxes.
  • 10.1.1 Existing XAI Methods.: Existing graph-XAI methods use gradients or features, perturbations, decompositions, and surrogates, but 3D and equivariant geometric models remain underexplored.A 3D perturbation method learns point-wise noise importance, yet it does not account for invariance or equivariance of explanations.
  • 10.1.2 Potential Application Scenarios.: XAI can improve trustworthiness, reveal scientific knowledge, diagnose violations of physical laws, and guide drug, material, and protein design.Applications include validating molecular features, protein residues, conservation laws, fermion antisymmetry, binding sites, and property-relevant substructures.
  • 10.2 Out-of-Distribution Generalization: OOD shifts substantially degrade scientific AI performance, arising across quantum, molecular, protein, material, chemical-interaction, and PDE tasks.Shifts include system size, geometry, molecular scaffolds, unseen protein structures, material compositions, novel bindings, and PDE conditions or resolutions.
  • 10.2.1 Background and Settings.: Domain adaptation aligns labeled source and unlabeled or labeled target distributions, whereas domain generalization predicts unseen domains without pre-collected target samples.Causality and invariant learning provide theoretical foundations by modeling interventions and seeking stable predictive mechanisms across environments.
  • 10.2.2 OOD in AI for Quantum Mechanics.: OOD methods address scientific shifts through invariant representations, domain knowledge, self-supervision, ensembles, detection, and generation; NCLaw generalizes across several PDE shifts after one trajectory.NCLaw handles new geometries, initial and boundary conditions, temporal ranges, and multiphysics systems, while benchmarks such as GOOD, DrugOOD, and CardioTox support evaluation.
  • 10.2.10 Open Research Directions.: Future OOD research should identify causal factors, incorporate physical principles, and exploit known shift structure, including SE(3) equivariance for orientation changes.Scientific OOD remains broadly underexplored despite its importance for preventing failures in real-world applications.
  • 10.3 Foundation and Large Language Models: Self-supervised learning extracts labels from data, while foundation models transfer pretrained representations across tasks with limited labels; LLMs additionally support scientific reasoning and generative discovery.Contrastive methods use paired positives and negatives, whereas predictive methods generate targets from data subsets; molecular SSL incorporates motifs, atom-bond associations, and reaction context.
  • 10.3.1 Self-Supervised Learning.: Physics-informed self-supervision reduces expensive PDE-solver data requirements, while physics-informed DeepONet generalizes across PDE families and demonstrated out-of-distribution performance.PINNs optimize losses derived from governing equations; physics-informed DeepONet further outperformed its supervised counterpart.
  • 10.3.2 Single-Modal Foundation Models.: Foundation models advance protein structure prediction, generation, and interaction modeling, while molecule models span graph, 3D-coordinate, and SMILES representations without a dominant non-language-based model.Protein systems include AlphaFold, RoseTTAFold, RFdiffusion, Chroma, and multimer extensions; molecular string generation motivates validity-preserving representations such as SELFIES.
  • 10.3.3 Natural Language-Guided Scientific Discovery.: Language-guided science uses multimodal alignment for high-level control, with bi-encoders supporting efficient cross-modal retrieval and joint encoders supporting finer interactions and generation.Scientific LLM adaptation requires choosing modalities, formulating sequential tasks, and selecting among training from scratch, fine-tuning, or few-shot prompting, often with external knowledge and tools.
  • 10.3.4 Open Research Directions.: Open directions include acquiring diverse, less noisy scientific data, designing multimodal and distribution-robust algorithms, and extending SSL and foundation models to underexplored quantum and scientific domains.Quantum SSL could learn symmetry and physical rules from unlabeled data, while text-based foundation models may transfer knowledge across systems and modalities.
  • 10.4 Uncertainty Quantification: Reliable uncertainty quantification is essential for robust decisions in physics-constrained prediction and generation, but scalable and efficient Bayesian or variational methods remain needed.The challenge spans neural ODEs, DeepONets, and other predictive or inverse-generative models under data and model uncertainty.
  • 10.4.1 Uncertainty Quantification: Introduction and Background.: Uncertainty arises from stochastic processes, incomplete or misspecified models, nonstationary dynamics, noisy data, and numerical error, and separates into irreducible aleatoric versus reducible epistemic components.Quantifying these sources supports safer modeling, active learning, Bayesian optimization, and discovery of new materials and compounds.
  • 10.4.2 Uncertainty Quantification in Computational Science.: Computational UQ distinguishes forward propagation of input randomness from inverse model calibration, using methods such as Monte Carlo, surrogate approximations, and Bayesian posterior inference.Bayesian approaches treat parameters or inputs as random variables and update prior beliefs using observed outputs, whereas alternatives include interval, confidence-score, distance-based, and UQ4K methods.
  • 10.4.3 Uncertainty Quantification in Deep Learning.: Deep-learning UQ represents aleatoric and epistemic uncertainty through Bayesian neural networks, approximate inference, ensembles, conformal prediction, evidential learning, and graph-specific models.Evidential deep learning explicitly models lack of evidence and has been useful for out-of-distribution detection, while graph methods model weight, topology, or prediction uncertainty.
  • 10.4.4 Uncertainty Quantification in AI for Science.: Scientific-AI studies show UQ can identify inaccurate or problematic predictions across molecular, quantum, and PDE tasks, while performance varies with the task and evaluation setup.Examples include uncertainty-informed data-quality assessment, bias mitigation and active learning, electron-density error detection, and PDE OOD detection.
  • 10.4.4 Uncertainty Quantification in AI for Science.: Scientific AI UQ can support OOD detection, noise identification, active learning, and Bayesian experimental design without harming predictive performance, but no method consistently dominates across tasks and setups.Benchmarks report task-dependent winners, including ensembles, random forests, and Gaussian processes, while other studies find no consistently superior method.
  • 10.4.5 Open Research Directions.: Open directions include high-fidelity benchmarks and standardized metrics, UQ methods incorporating physical or biological knowledge, and scalable approaches that balance uncertainty quality with computational cost for large models.Existing methods are either simplistic or restrictive, such as MC dropout, or expensive, such as deep ensembles and Bayesian neural networks; subspace-based approximations may improve scalability.

11 LEARNING, EDUCATION, AND BEYOND

This section surveys resources for learning AI and science, identifies paradigm shifts toward interdisciplinary knowledge, diverse communities, and richer education, and calls for a unified roadmap and collective action.

  • 11.1 Existing Resources for Fundamental AI and Science: Books, courses, libraries, and conferences provide foundational resources across individual AI and scientific fields, with representative compilations in Tables 35 and 37.These materials support knowledge acquisition, collaboration, and computational practice, although the listed resources are not exhaustive.
  • 11.2 Paradigm Shifts in AI for Science: AI for Science must transcend disciplinary boundaries by integrating perspectives and methods from fields including physics, biology, chemistry, and artificial intelligence.Examples include physics-informed neural-network interpretation [Sorscher et al. 2022; Di Giovanni et al. 2023], generative modeling with dynamical systems and control [Song et al. 2020; Xu et al. 2022a; Liu et al. 2022a; Berner et al. 2022], and protein structure prediction.
  • 11.2 Paradigm Shifts in AI for Science: A diverse, flexible community and expanding educational ecosystem are emerging through workshops, symposiums, summer schools, tutorials, reading groups, and academia–industry collaboration.AI-related articles in ACS Omega increased 133% from 2020 to 2021, while systematic curricula remain difficult because methods advance rapidly and span disciplines.
  • 11.3 Prospective and Proposed Actions: Because resources remain fragmented and lack a cohesive roadmap, the field should develop unified educational materials, collaborative platforms, and effective knowledge-collection methods.Researchers, practitioners, students, and contributors to AI for Science subareas are encouraged to participate in building this shared system.

12 CONCLUSION

This work presents a technical, unified review of AI for science across spatial and temporal scales, covering problem setups, challenges, approaches, datasets, benchmarks, and future directions. It also provides categorized resources for learning while acknowledging that the evolving field requires continued expansion.

  • Review scope: The review organizes AI-for-science research by the spatial and temporal scales of physical-world modeling, with precise problem setups, key challenges, major approaches, datasets, benchmarks, and future directions.It addresses challenges including symmetry, interpretability, out-of-distribution generalization, causality, and uncertainty quantification.
  • Resources and outlook: The work compiles categorized resources to facilitate learning and education in the emerging field of AI for science.The authors emphasize that the review is neither comprehensive nor conclusive and will evolve with the field and community feedback.

A CLASSIFYING AND COMPUTING IRREDUCIBLE REPRESENTATIONS

The section explains how Schur’s lemma supports computing irreducible representations (irreps), while characters and regular representations classify finite-group irreps. For semisimple Lie algebras, Cartan subalgebras and highest weights provide the corresponding classification framework.

  • Computing irreducible representations: Schur’s lemma turns decomposition of a reducible representation into a nullspace and eigenspace computation, yielding irreducible components without prior knowledge of the irreps.Commuting linear maps are found by solving Qρ_X(g)−ρ_X(g)Q=0; random linear combinations of a nullspace basis have eigenspaces that, with high probability, decompose the representation.
  • Classifying irreps for Finite Groups: For finite groups, characters identify isomorphic representations through conjugacy-class traces, while decomposing the regular representation recovers all possible irreducible representations.The regular representation can be constructed from the group multiplication table and decomposed into irreps.
  • Classifying irreps for the Semisimple Lie Groups: For semisimple Lie algebras, complete reducibility and Lie-algebra Schur’s lemma extend the decomposition procedure, while Cartan subalgebras split representations into weight spaces.Weights are simultaneous eigenvalue functionals on a Cartan subalgebra; roots arise as nonzero weights of the adjoint representation, and the Killing form supplies an inner product.
  • Classifying irreps for the Semisimple Lie Groups: The highest-weight theorem states that irreducible representations have unique dominant integral highest weights, equal highest weights imply isomorphism, and every dominant integral weight occurs.Thus dominant integral weights classify all irreps; for 𝔰𝔬(3), they are indexed by nonnegative half-integers, but only integer labels correspond to representations of SO(3).
Loading 2307.08423v6…