Source-linked AI summary

Machine learning for electronically excited states of molecules

Julia Westermayr, Philipp Marquetand

arXiv:2007.05320v1physics.chem-phstat.ML

TL;DR

Accurate excited-state simulations are important across photochemistry and related fields but remain computationally expensive and methodologically difficult. This review surveys how ML can accelerate and extend excited-state quantum chemistry, covering static properties, dynamics, training data, and methodological pitfalls. The reviewed studies improve many stages of excited-state calculations, while transferable models and ML representations of multireference wave functions remain open goals.

  • Problem

    Accurate theoretical treatment of molecular excited states is computationally expensive, complex, and often requires expert knowledge.

  • Method

    The review synthesizes ML applications for excited-state quantum chemistry alongside quantum-chemical methods, nonadiabatic dynamics, training-set generation, and associated challenges.

  • Results

    Reviewed ML studies successfully improve almost all aspects of excited-state quantum chemistry, including active-space selection, prediction of quantum-chemical outputs, and interpretation of theoretical studies.

  • Takeaways & Limitations

    ML may enable more efficient and flexible excited-state simulations and support future prediction of photochemistry in larger molecular systems.

  • Takeaways & Limitations

    Transferable excited-state models remain limited, and fitting regions near conical intersections can be incomplete because singularities cause algorithmic failures.

Abstract

from arXiv · show

Electronically excited states of molecules are at the heart of photochemistry, photophysics, as well as photobiology and also play a role in material science. Their theoretical description requires highly accurate quantum chemical calculations, which are computationally expensive. In this review, we focus on how machine learning is employed not only to speed up such excited-state simulations but also how this branch of artificial intelligence can be used to advance this exciting research field in all its aspects. Discussed applications of machine learning for excited states include excited-state dynamics simulations, static calculations of absorption spectra, as well as many others. In order to put these studies into context, we discuss the promises and pitfalls of the involved machine learning techniques. Since the latter are mostly based on quantum chemistry calculations, we also provide a short introduction into excited-state electronic structure methods, approaches for nonadiabatic dynamics simulations and describe tricks and problems when using them in machine learning for excited states of molecules.

1 Introduction

Machine learning is emerging in excited-state quantum chemistry because accurate simulations are computationally expensive and difficult, while applications span the calculation pipeline from inputs to analyzed outputs. This review surveys these applications, their methods, training-set challenges, and limitations.

  • Motivation: ML for quantum chemistry is developing rapidly, but excited-state applications remain comparatively young because reference calculations and models are especially complex and expensive.The review describes this field as still being in its initial stage.
  • Scope: The review focuses mainly on molecular singlet and triplet states, while excluding other excitation types and treating condensed-phase excited states as challenging for conventional approaches.Most existing excited-state ML studies concern singlet states, with fewer studies addressing triplet states.
  • Motivation: Excited-state quantum chemistry is crucial for understanding photochemical and biological processes, but experiments alone cannot directly describe their exact electronic mechanisms.Theory can complement spectroscopy and clarify photo-induced reaction mechanisms.
  • Scope: ML can target excited-state quantum chemistry at many stages, including active-space selection, quantum-chemical calculations, secondary outputs, tertiary properties, and analysis.Examples include energies, derivatives, spectra, excitation energies, and structure–property correlations.
  • Challenges: The authors emphasize training-set generation and quantum-chemical reference methods as bottlenecks that constrain broader application of ML to excited states.The review covers theoretical methods, molecular-dynamics applications, training-set construction, and unresolved methodological problems.
  • Review structure: The review organizes state-of-the-art approaches by single- versus multi-state and single- versus multi-property models, with emphasis on ML force fields for excited states.It also discusses how ML may improve simulations and identifies future research directions.

2 General Background: From the Ground State to the Excited States

Excited-state chemistry differs fundamentally from ground-state modeling because molecules may access multiple spin manifolds, minima, transition regions, and crossings. These features make accurate simulations and transferable force fields difficult, although ML can combine quantum-chemical accuracy with force-field efficiency.

  • Excited-state processes: Excitation can populate a manifold of electronic states, followed by internal conversion, intersystem crossing, fluorescence, phosphorescence, or chemical reactions.The accessible states depend on photon energy and transition properties such as oscillator strength and transition dipole moments.
  • Ground-state modeling: Conventional force fields enable large-system and nanosecond-scale ground-state simulations but have limited accuracy and cannot describe bond formation and breaking reliably.Reactive force fields exist but still face generally low accuracy and have not become mainstream.
  • Role of ML: ML can combine ab-initio accuracy with force-field efficiency when trained on comprehensive reference data containing energies, forces, and ground-state properties.The stated benefit is retaining reference-method accuracy while making inferences much faster.
  • Excited-state surfaces: Excited-state potential-energy surfaces contain multiple singlet and triplet states, local minima, transition regions, and crossings, making separate-state treatment inaccurate.This complexity restricts straightforward transfer of ground-state quantum-chemical and ML approaches.
  • Computational scaling: As molecular size increases, electronic states become closer in energy and more states must be considered, restricting accurate excited-state methods to systems with only a few dozen atoms.The increasing number of states further raises computational expense.
  • Transferability: Transferable excited-state ML force fields remain unavailable across chemical compound space, with accurate demonstrations limited to about 20 atoms and 3 electronic states of distinct multiplicity.The review identifies larger, chemically diverse systems such as proteins and DNA as an important target for future development.

3 Quantum Chemical Theory and Methods

The review introduces the quantum-theoretical foundations and notation needed to interpret excited-state calculations and their use as ML training data. It emphasizes the distinction between electronic and nuclear motion and the challenges of treating multiple excited states.

  • Purpose of the background: Excited-state quantum-chemical calculations provide the training data for ML models, so their theoretical foundations and nomenclature are essential for evaluating ML applications.The review uses this background to assess reference methods and applications such as excited-state molecular dynamics.
  • Excited-state challenges: The review explains differences between ground- and excited-state computations, particularly those arising from the need to treat a manifold of excited states.These differences also identify problems that can affect ML models.
  • Notation: Notation varies across the literature, including different uses of terms such as nonadiabatic couplings and interstate couplings.The review adopts consistent notation while acknowledging these terminology differences.
  • Theoretical foundation: The theoretical framework begins with separation of electronic and nuclear degrees of freedom through the Born–Oppenheimer approximation, which is partly lifted in nonadiabatic dynamics.Nonadiabatic simulations therefore include electron–nuclear coupling.

3.1 Electronic Structure Theory for Excited States

Excited-state electronic structure methods approximate the electronic Schrödinger equation through wave-function or density-based formulations, balancing accuracy, computational cost, and practical feasibility. Wave-function methods expand electronic states with determinants or active-space configurations, while density-functional approaches depend critically on the chosen functional and can fail for important excited-state regimes.

  • Electronic structure calculations seek molecular potential energies and properties using Wave Function Theory or Density Functional Theory.
  • The electronic Schrödinger equation defines electronic states through an N-electron wave function, but exact solutions become infeasible beyond very small molecules.Approximated wave functions are therefore introduced for larger systems.
  • Density Functional Theory: Density Functional Theory is computationally efficient but lacks a universal, systematically improvable functional, and its accuracy depends on the targeted excitation regime.Functional choices differ for valence, Rydberg, vertical, and long-range charge-transfer excitations.
  • Wave Function Theory (WFT): Configuration Interaction expands a Hartree-Fock reference with Slater determinants, while truncated variants such as CIS and CISD include selected single and double excitations.Full-CI considers all possible configurations but is infeasible for almost all molecular systems more complex than helium.
  • Wave Function Theory (WFT): Multi-reference methods such as CASSCF use manually selected active spaces and perform a Full-CI calculation within them to describe electronically complex states.The active space separates inactive doubly occupied, active, and inactive empty orbitals.
  • Wave Function Theory (WFT): Multi-reference approaches are difficult to use for excited-state simulations because active-space choices, intruder states, numerical instabilities, and algebraic complexity can produce inconsistent or infeasible calculations.Training-set generation for machine learning can become infeasible when multi-reference calculations require millions of configuration state functions.
  • Density Functional Theory: Linear-response TDDFT is often used for medium-sized and large systems because of its efficiency, but standard implementations describe conical intersections with incorrect dimensionality.CI-corrected and hole-hole Tamm-Dancoff variants can recover missing couplings and correct this dimensionality.

3.2 Bases

Excited-state potential energy surfaces can be expressed in several bases connected by unitary transformations, each emphasizing different state properties and couplings. Diabatic potentials are smooth and therefore particularly favorable for numerical applications, including machine learning, whereas adiabatic surfaces can become non-smooth near conical intersections.

  • Basis types: The review distinguishes diabatic, adiabatic, and spin-adiabatic bases, whose common names and relationships are summarized in Table 1 and Figure 6.These bases arise from different expansions of the total molecular wave function and are connected by unitary transformations.
  • Adiabatic basis: Adiabatic potential energy surfaces within one spin multiplicity are energy-ordered but become non-smooth near conical intersections and avoided crossings.Non-adiabatic coupling norms can show sharp spikes at avoided crossings and nearly vanish elsewhere.
  • Spin-diabatic basis: When multiple spin multiplicities are included, singlets and triplets are adiabatic within their own multiplicities but diabatic with respect to each other.Spin-orbit couplings connect states of different multiplicities, while triplet components associated with different magnetic quantum numbers are degenerate.
  • Diabatic basis: Strictly diabatic bases do not exist in practice for polyatomic systems, so approximate quasi-diabatic surfaces are fitted and depend on the method and reference point.The resulting diabatic potentials are not unique.
  • Diabatic basis: Diabatic potentials cross at avoided crossings, preserve electronic character, and remain smooth as functions of nuclear coordinates.Their smoothness makes diabatic potential energy surfaces highly favorable for numerical applications, including machine learning.
  • Basis transformations: The MCH and diabatic bases are interconverted by a unitary matrix that is determined up to an arbitrary sign from the wave-function phase.For two states, the transformation matrix is a rotation matrix.

3.3 Excited-State Dynamics Simulations

Excited-state dynamics simulations solve time evolution by repeatedly obtaining electronic potentials and using them to propagate nuclear motion. Quantum approaches require global potential representations, whereas classical trajectory methods can evaluate potentials on the fly and are more suitable for larger systems, although coupling calculations remain costly.

  • General framework: Each excited-state dynamics time step solves the electronic problem to obtain potentials and forces, then propagates the nuclear equations of motion.This procedure follows the time-dependent Schrödinger-equation framework for isolated molecular systems.
  • Quantum approaches: Quantum nuclear dynamics uses wave functions spread across nuclear coordinates, requiring potentials computed in advance and interpolated or stored globally.Machine learning is intended to improve interpolation of these potential energy surfaces.
  • Classical approaches: Classical and mixed quantum-classical dynamics propagate trajectories at individual geometries, permitting on-the-fly potential calculations when fewer geometries are visited than a global representation requires.Trajectory surface hopping is the most popular mixed quantum-classical method; each trajectory has one active state but can transition between states.
  • Quantum approaches: Exact quantum nuclear dynamics scales exponentially with nuclear degrees of freedom and is typically limited to systems containing fewer than 5 atoms.Even for such small systems, calculating the potential energy surfaces can justify machine-learning approximations.
  • Approximations: MCTDH and related quantum or semiclassical methods extend accessible system sizes or include quantum effects, but computational costs remain substantial and reduced-dimensionality choices can affect comparisons with full-dimensional classical dynamics.The relative quality of reduced-dimensionality quantum dynamics and full-dimensional classical dynamics remains system-dependent.
  • Approximations: Ring-polymer dynamics includes nuclear quantum effects but suffers high computational effort because many replicas are required.Semiclassical methods can accurately treat systems up to tens of atoms, while larger systems are dominated by cheaper mixed quantum-classical methods.
  • Surface hopping: Surface-hopping transition probabilities may use energy differences, non-adiabatic couplings, or arbitrary couplings, but methods requiring non-adiabatic couplings face a major computational bottleneck.Coupling calculations remain among the most expensive parts of quantum-chemical calculations.

3.4 Dipole Moments and Spectra

Dipole moments connect excited-state calculations with spectroscopic observables, especially absorption spectra. Machine-learning models commonly target relative transition properties or phase-independent quantities because absolute transition dipoles have sign ambiguities and can be difficult to reproduce.

  • Permanent dipoles: Permanent dipole moments are important for comparing theory with experiment and can be used to compute infrared spectra from molecular-dynamics simulations.The infrared spectrum is obtained from the Fourier transform of a time autocorrelation function involving the dipole-moment derivative.
  • Transition dipoles: Excited-state simulations commonly use transition dipole moments, while ground- and excited-state permanent dipoles can differ strongly after light excitation.The difference reflects frequency shifts and altered electron distributions.
  • Machine-learning targets: A latent-charge neural-network model can fit transition and permanent dipoles by inferring point charges rather than learning them directly.The model uses atomic position vectors in constructing the dipole.
  • Absorption spectra: Many computational studies fit relative transition-dipole values instead of absolute values to obtain reasonably accurate absorption spectra for comparison with experiments.Absorption spectra can be calculated from excited-state energies and oscillator strengths, which are proportional to squared transition dipole moments.
  • Machine-learning targets: Oscillator strengths or dipole-vector lengths can avoid the arbitrary-sign problem of transition dipole moments caused by the wave-function phase.The transition dipole moment itself is not uniquely defined with respect to sign.

4 Data Sets for Excited States

Excited-state ML requires training data that accurately covers molecular configurations, electronic states, derivatives, and couplings despite expensive reference calculations, noisy outputs, and arbitrary wave-function phases. Sampling strategies, phase corrections, and iterative or adaptive learning can reduce data-generation demands, but conical intersections and high-density state manifolds remain difficult.

  • Reference data: Excited-state PES generation is costly because accurate methods, forces, couplings, and many electronic states may all be required.The computational burden increases substantially for systems with dense electronic states.
  • Reference data: Reference-method selection must match the target properties, number of excited states, system size, and complexity of the photochemical process.Couplings require particular care because not all quantum-chemical methods provide them and their signs can jump along reaction coordinates.
  • Reference data: Multi-reference training sets can become infeasible; one example required 19,302,445 configuration state functions for 356 cyclopentoxy data points.For such systems, generating an ample training set may be too expensive even before ML training begins.
  • Phase correction: Transition dipole moments, NACs, and SOCs have arbitrary signs because they arise from pairs of electronic states, complicating conventional ML fitting.Phase correction or phase-free training addresses this issue; internal phase correction was implemented in SchNarc for photodynamics simulations.
  • Limitations: Conical-intersection singularities can make phase-correction algorithms fail, forcing threshold-based data removal and leaving those PES regions less comprehensively fitted.Manual sign fitting is also reported as tedious for larger systems and higher-dimensional descriptions.
  • Sampling strategies: Training on 1,000 data points reproduced reference dynamics, while structure-based sampling reduced static-calculation data requirements by up to 90% for methyl chloride.The broader sampling procedure first generated 10,000 points from low-lying PES regions using an inexpensive method.
  • Limitations: Equilibrium-structure training sets may be poorly suited to photodynamics because excited-state motion rapidly reaches conformations far beyond the sampled region.This motivates sampling procedures designed specifically for excited-state trajectories.

5 ML Models

ML models for excited-state quantum chemistry use regressors and molecular descriptors to predict quantum-chemical outputs, with kernels and neural networks as prominent approaches. Their accuracy and scalability depend strongly on training data, representation, and model structure.

  • General considerations: Training-set quality, regressor choice, and molecular descriptors jointly constrain the attainable accuracy of ML models.Improper regressors or descriptors can produce inaccurate predictions.
  • Regression models: Regression relates molecular inputs X to quantum-chemical outputs Y, while linear regression provides a baseline for minimum achievable accuracy.Training optimizes weights and biases by minimizing a loss such as MAE- or MSE-based objectives.
  • Kernel methods: Kernel methods such as KRR and GPR use similarity measures and nonlinear kernels to extend ridge regression beyond linear relationships.Kernel functions measure distances between a query compound and training compounds, while regularization helps prevent overfitting.
  • Kernel methods: Kernel methods have few hyperparameters and can provide nearly exact solutions, but kernel-matrix inversion becomes costly as training sets grow.Their memory requirements can make inversion infeasible on current computers.
  • Kernel methods: Standard kernel implementations generally map one input to one output, so multiple excited states commonly require separate models.Forces can be described for the ground state or a single excited state using approaches including KRR, sGDML, or SOAP-based GPR.
  • Neural networks: Neural networks are flexible parametric functions that can fit large datasets and map one molecular input to multiple quantum-chemical outputs.For excited-state manifolds, a network can output several states and additional properties, with forces obtained as derivatives of NN potentials.
  • Descriptors and features: Molecule-wise descriptors are easy and inexpensive for small systems but may omit angular information, whereas atom-wise representations support arbitrary system sizes yet remain insufficiently validated for excited-state surfaces.Current atom-wise excited-state fits have focused on small molecules, and accurate medium- or large-system excited-state PESs remain out of reach.

6 Application of ML for Excited States

The review organizes ML studies of excited states and their properties around applications that improve static and dynamics calculations. It emphasizes regressors, descriptors, training sets, and predicted properties as the basis for comparing approaches.

  • Scope and organization: The review classifies ML studies of excited states according to their applications in static and dynamics calculations.It focuses on the regressor, descriptor, training set, and property used in each approach.

6.1 Parameters for Quantum Chemistry

ML can assist quantum-chemistry parameter selection by choosing active spaces for multireference calculations. This addresses the manual selection of active orbitals and electrons, although choosing between multireference and single-reference methods remains unresolved.

  • Active-space selection: An XGBoost classification protocol automatically selects relevant active spaces for multireference methods in molecular systems.The approach was demonstrated for diatomic molecules in the dissociation limit using molecular-orbital bond order and average electronegativity.
  • Active-space selection: The protocol can avoid the tedious manual selection of active orbitals and active electrons.However, users must still decide whether a multireference or single-reference method is appropriate.

6.2 ML of Primary Outputs

For excited states, ML models providing primary quantum-chemistry outputs remain unavailable according to the review. Directly targeting wave functions or density functionals is difficult, but could enable wave-function analysis and additional excited-state insights.

  • Primary outputs: No ML models for providing primary quantum-chemistry outputs for excited states were known to the authors.Primary outputs include the N-electron wave function and ML density functionals.
  • Primary outputs: A successful ML approach to excited-state primary outputs could support wave-function analysis and provide additional insights.The passage presents this as a potential benefit rather than an achieved result.

6.3 ML of Secondary Outputs

Machine learning models fit secondary quantum-chemical outputs—including potential-energy surfaces, couplings, and dipole moments—in single-state or multi-state forms for excited-state dynamics. These models can reproduce reference photodynamics while substantially reducing computational cost, although accuracy in critical regions and the treatment of diabatic surfaces remain important constraints.

  • Scope: Secondary-output models learn potential-energy surfaces, spin–orbit couplings, nonadiabatic couplings, and transition or permanent dipole moments in adiabatic or diabatic representations.These quantities can be fitted using single-state or multi-state approaches.
  • Diabatic representations: Diabatic potentials are smooth and well matched to machine-learning models, but generating them remains laborious, motivating ML-assisted diabatization.The limiting step is the tedious procedure required to generate diabatic potential-energy surfaces.
  • Applications: Neural-network and kernel models supported excited-state dynamics, including absorption, population transfer, and full-dimensional simulations with fitted potentials and couplings.Applications included NH3, H2O, formaldehyde, SO2, CSH2, and other molecular systems.
  • Computational acceleration: A 1 ns photodynamics simulation required approximately two months with ML models, compared with an estimated 19 years using MR-CISD directly.The dynamics used hopping probabilities based on ML-fitted nonadiabatic couplings, and neural networks replaced the reference method during dynamics.
  • Reliability: Small ML errors in critical potential-energy-surface regions can produce completely wrong photodynamics, despite faithful reproduction of multireference potential-energy curves elsewhere.This highlights that curve-fitting accuracy alone is insufficient for reliable dynamics.
  • Computational acceleration: For CH2NH+2, learning nonadiabatic couplings directly was faster than approximating them from Hessians, whose second-order derivatives reduced efficiency by about tenfold.GPU acceleration improved ML-PES Hessian calculations by about 5–10 times, depending on the molecule and GPU.
  • Computational acceleration: Cheaper reference methods do not necessarily yield comparable speedups: CSH2 showed less acceleration, while slow population transfer reduced the difference between learned and approximated couplings.Slow population transfer required fewer Hessian evaluations for estimating hopping probabilities.

6.4 ML of Tertiary Outputs

Machine learning extends excited-state predictions from energies and dipoles to spectra, spin ordering, and electronic-structure-driven molecular design. The reviewed studies report accurate or qualitatively useful predictions across UV, X-ray, and experimental-spectrum analysis tasks, with descriptor and state encoding choices affecting performance.

  • Spectra: Excited-state energies and dipole moments can be combined to compute oscillator strengths and energy gaps for modeling UV absorption spectra.These secondary outputs provide inputs for tertiary spectral properties.
  • UV spectra: Delta learning with KRR reached CC2 accuracy for the two lowest excited singlet energies and corresponding oscillator strengths in the QM8496 database.The database contained approximately 20,000 organic molecules.
  • Descriptor design: Transition-energy prediction from QM9 was not sufficiently accurate with atom-wise descriptors, motivating advanced non-local descriptors and explicit encoding of electronic-state information.Atom-wise descriptors performed well for HOMO–LUMO gaps but less well for transition energies.
  • UV spectra: A random forest trained on 500,000 PubChemQC molecules reportedly outperformed previous models for oscillator strengths and excitation energies of the most probable organic-molecule transition.The analysis identified nitrogen-containing heterocycles as important for high oscillator strengths and suggested applications to fluorophore design.
  • Spectra: Convolutional neural networks with Coulomb matrices and DTNNs outperformed simpler neural networks while achieving good agreement with reference DFT spectra.The models predicted molecular spectra using orbital energies and Gaussian broadening with a full width at half maximum of 0.5 eV.
  • X-ray spectra: Neural networks predicted Fe K-edge X-ray near-edge spectra from local radial distributions using 9,040 training data points.The inputs were generated from arbitrary systems in the Materials Project Database.
  • Experimental-spectrum analysis: Unsupervised clustering and Gaussian-process regression produced fingerprint spectra that differentiated functionalized amorphous surfaces and approximated experimental sample composition semiquantitatively.The spectra combined electronic-structure and SOAP-based structure information.
  • Spin states: Deep neural networks assigned the correct spin in most tested spin-crossover complexes using features describing ligand–metal bonding, metal identity, oxidation state, charge, and denticity.Spin-state ordering helps evaluate catalytic and material properties of metal complexes.

6.5 ML-Assisted Analysis

Machine learning can also analyze outputs from excited-state simulations and experimental spectroscopy rather than only predict quantum-chemical quantities. The reviewed applications include dissociation-time prediction, photoluminescence analysis, transient-absorption analysis, and correlations between spectra and electronic coupling.

  • Simulation analysis: The low cost of ML-enabled dynamics and spectral prediction permits many trajectories and long simulations, but analyzing the resulting production data can become the time-limiting step.This issue was identified in work on dissociation times of 1,2-dioxetane.
  • Spectroscopy: LumiML uses linear regression on computer-generated photoluminescence data to predict decay-rate distributions of perovskite nanocrystals.The software was applied to data from femtosecond broadband fluorescence upconversion spectroscopy and can also analyze transient absorption spectra.
  • Spectroscopy: Bayesian neural networks were used to relate nanoaggregate structure to electronic coupling in semiconducting materials through absorption spectra.The application connects experimental-spectrum analysis with correlations relevant to material systems.

7 Conclusion and Future Perspectives

The review finds that machine learning now improves many aspects of excited-state quantum chemistry, but current approaches remain molecule-specific, data-intensive, and dependent on human intervention. Future progress requires transferable representations and stronger integration with accurate electronic-structure methods.

  • Reference methods: About half of the reviewed training sets use multi-reference methods, rising to about 70% among studies targeting excited-state dynamics.Dynamics studies use single-reference methods in 15% of cases, with another 15% using model Hamiltonians or analytical potentials.
  • Data requirements: Photodynamics remains data-intensive: many thousands of points are needed for a few excited-state potentials, and full-dimensional dynamics beyond 12 atoms has not yet been investigated.Active learning, adaptive sampling, and structure-based sampling are identified as important for generating meaningful dynamics training sets.
  • Method limitations: Single-reference DFT lowers training-set cost but can produce qualitatively incorrect potential-energy surfaces in regions such as dissociation, leaving dynamics training sets incomplete if those regions are excluded.The review identifies Δ-learning and transfer learning as possible ways to address this problem.
  • Data preparation: Training excited-state properties requires phase correction because coupling values and related properties have arbitrary signs between electronic states.Phase correction has been applied to coupling values and dipole moments before conventional ML fitting.
  • ML models: About two thirds of the reviewed studies rely on neural networks, while kernel ridge regression is mainly used for diabatic-potential interpolation and multi-molecule studies.Only a few studies address extrapolation across chemical compound space for excited states.
  • Transferability: Current models are often molecule-specific because descriptors usually represent whole molecules, hindering transfer across chemical compound space.Transferable models require descriptors that encode atomic chemical and structural environments and support molecules of arbitrary size and composition.
  • Scope and contributions: ML studies address active-space selection, quantum-chemical outputs, excited-state dynamics, spectra, and interpretation of theoretical results.The reviewed applications span secondary and tertiary outputs as well as analysis of computed and experimental data.
  • Future perspectives: The review states that ML is far from replacing quantum-chemical methods or becoming routine, and that human intervention remains necessary, especially beyond isolated molecules.The authors frame ML primarily as a means to improve existing methods rather than replace them completely.
Loading 2007.05320v1…