Source-linked AI summary
ANI-1: An extensible neural network potential with DFT accuracy at force field computational cost
Justin S. Smith, Olexandr Isayev, Adrian E. Roitberg
TL;DR
The paper addresses the lack of a transferable, accurate potential for complex organic molecules. It develops ANI using modified symmetry-function representations and reports extensibility to larger molecules than those in training, with accuracy dependent on the training data.
Problem
Existing neural-network potentials had not demonstrated true transferability across complex organic chemical environments.
Method
ANI uses modified Behler–Parrinello symmetry functions to construct single-atom atomic environment vectors for neural-network potentials.
Results
ANI potentials showed extensibility to larger molecules than those included in the training data set, with extensibility increasing with data-set size without plateauing at the current size.
Takeaways & Limitations
ANI provides a transferable neural-network potential for organic molecules within its trained chemical domain.
Takeaways & Limitations
The accuracy of ANI depends entirely on the data used during training.
Abstract
from arXiv · showhide
Deep learning is revolutionizing many areas of science and technology, especially image, text and speech recognition. In this paper, we demonstrate how a deep neural network (NN) trained on quantum mechanical (QM) DFT calculations can learn an accurate and fully transferable potential for organic molecules. We introduce ANAKIN-ME (Accurate NeurAl networK engINe for Molecular Energies) or ANI in short. ANI is a new method and procedure for training neural network potentials that utilizes a highly modified version of the Behler and Parrinello symmetry functions to build single-atom atomic environment vectors as a molecular representation. We utilize ANI to build a potential called ANI-1, which was trained on a subset of the GDB databases with up to 8 heavy atoms to predict total energies for organic molecules containing four atom types: H, C, N, and O. To obtain an accelerated but physically relevant sampling of molecular potential surfaces, we also propose a Normal Mode Sampling (NMS) method for generating molecular configurations. Through a series of case studies, we show that ANI-1 is chemically accurate compared to reference DFT calculations on much larger molecular systems (up to 54 atoms) than those included in the training data set, with root mean square errors as low as 0.56 kcal/mol.
1 Introduction
Large-molecule energetics requires methods that balance quantum-mechanical accuracy with force-field-scale cost, while existing empirical and neural potentials face transferability limits. The paper introduces ANI as a transferable deep-learning potential for diverse organic molecules.
- Motivation: Classical force fields enable large-scale simulations at reduced computational cost but are generally reliable only near equilibrium.Their typical nonreactive formulation limits applicability to chemical reactions and transition states.
- Existing limitations: ReaxFF can study chemical reactions and transition states, but generally requires reparametrization from system to system.This limits its out-of-the-box transferability and makes system-specific benchmarking usually unavoidable.
- Representation requirements: Transferable molecular representations must be rotationally and translationally invariant, permutation-invariant for identical atoms, and uniquely describe molecular conformation.The original symmetry functions were limited by feature construction and atomic-number differentiation in complex chemical environments.
- Existing limitations: Existing neural-network potentials had not demonstrated true transferability between complex chemical environments in organic molecules.Prior work was limited to small test sets, equilibrium geometries, short molecular-dynamics trajectories, or selected systems.
- Contribution: ANI introduces modified symmetry functions that build single-atom atomic environment vectors to represent complex molecular systems.The method aims to learn statistically diverse molecular interactions across conformational and configurational space.
- Contribution: ANI potentials are designed to predict energies for molecules within an organic-molecule training domain and extend to larger molecules than those in training.The paper presents this extensibility as evidence of transferability beyond the training data set.
2 Theory and neural network potential design
ANI represents each atom’s local chemical environment with modified symmetry-function vectors, feeds these vectors into atom-specific neural networks, and sums atomic outputs to predict molecular energy. The design targets transferability across molecular structures while retaining near-linear computational scaling.
- Motivation: The model addresses standard neural-network-potential limitations involving data requirements, fixed input sizes, and transferability between molecules.The paper identifies increasing degrees of freedom, non-transferable input dimensions, and permutation invariance as central design challenges.
- Atomic representations: ANI uses modified Behler–Parrinello symmetry functions to construct atomic environment vectors encoding radial and angular local chemical features.Each vector is computed for an atom and supplied to a neural network associated with its atomic number.
- Energy model: The total molecular energy is computed as a sum of atomic contributions, enabling near-linear computational scaling with added cores or GPUs up to the system’s atom count.Each atomic environment vector is processed to produce an atomic energy contribution before summation.
- Symmetry-function design: Modified angular functions add arbitrary angular shifts and radial-shell resolution, producing more distinctive representations of bonding, ring, and functional-group patterns.The modifications reduce overlap between angular regions and help keep environment-vector elements smaller.
- Atomic-number specificity: ANI distinguishes atomic numbers through separate radial and angular components for element types and element pairs, improving training error and transferability on diverse multi-molecule sets.The original functions treated atoms identically in their summations, limiting atom-type discrimination.
- Sampling: Normal Mode Sampling is proposed to generate accelerated, physically relevant configurations for sampling molecular potential surfaces.The method perturbs equilibrium structures along calculated normal modes within a maximum-energy window.
3 Methods
The ANI-1 dataset combines near-equilibrium organic molecules with DFT reference energies and Normal Mode Sampling configurations. It covers H, C, N, and O molecules from GDB-11, producing millions of conformations for neural-network training and evaluation.
- Data scope: The study limits ANI-1 to organic molecules containing H, C, N, and O and restricts configurations to near-equilibrium conformations.The paper states that broader atom types and full potential-surface sampling would make the dataset impractical to construct.
- Molecule selection: ANI-1 is generated from a GDB-11 subset containing molecules with up to 8 C, N, and O atoms after fluorine-containing molecules are excluded.GDB-11 supplies molecular structures that are converted from SMILES strings to three-dimensional geometries.
- Sampling rationale: QM molecular dynamics is considered inefficient for sampling a large potential-surface energy window, while trajectory sampling can be biased toward a specific path.These limitations motivate the more stochastic Normal Mode Sampling approach.
- Configuration generation: Normal Mode Sampling perturbs optimized molecular structures along normal modes to produce configurations across a relevant energy window.The procedure calculates normal modes, generates K structures per molecule, and evaluates their single-point energies.
- Dataset scale: 17.2 million molecular energies are generated from approximately 58,000 small molecules.For each molecule, 80% of conformations are used for training and 10% each for validation and testing.
- Evaluation: The ANI-1 potential achieves training, validation, and test RMSE values of 1.2, 1.3, and 1.3 kcal/mol, respectively.The reported values compare ANI-1 predictions with DFT reference energies.
4 Results and discussion
ANI-1 is evaluated across statistical fitness, structural isomers, conformers, and potential-surface scans against DFT and semi-empirical baselines. Across these tests, increasing training-set size improves extensibility, while ANI-1 reproduces larger-molecule energies and energy differences with low errors and smooth surfaces.
- Statistical fitness: ANI-1 was trained on more than 56k small GDB-8 molecules using NMS-generated configurations spanning configurational and conformational space.The resulting data set contains over 80% of the 17.2 million ANI-1 data points.
- Statistical fitness: Extensibility increased with training-set size and showed no plateau at the current data-set size.Figure 3 compares training, validation, testing, and GDB-10 extensibility errors as training points increase.
- Statistical fitness: 1.8 kcal/mol RMSE was achieved for the full random-conformation test set, while adding more and more diverse training molecules improved transferability.The largest and most diverse training set produced the lowest reported RMSE in the training-size comparison.
- Structural and geometric isomers: 0.2 kcal/mol RMSE was achieved for C10H20 isomer energies, with ANI-1 correctly predicting the minimum structure, energy ordering, and ring-to-linear energy differences.DFTB and PM6 incorrectly predicted the relative stability of ring-containing structures and systematically underestimated some linear-isomer energies by about 6–7.
- Potential surface accuracy: ANI-1 reproduced smooth one-dimensional potential surfaces, with a 0.4 kcal/mol RMSE for one angle-bend scan and minima within 3.0° of DFT for a lisdexamfetamine dihedral scan.The semi-empirical methods underestimated dihedral barriers and produced unrealistic potential-surface shapes in the lisdexamfetamine case.
5 Conclusions
ANI-1 is presented as a transferable neural network potential for organic molecules, trained on small molecules yet evaluated on larger systems. It combines strong agreement with DFT, substantial speed advantages, and force-field-like scaling, while remaining dependent on training-data coverage.
- Conclusions: ANI-1 is presented as a transferable neural network potential for organic molecules trained on conformational and configurational data from molecules with up to 8 heavy atoms.The method is described as using a deep-learning architecture with substantial modifications to the Behler–Parrinello approach.
- Conclusions: ANI-1 was applied to larger systems of 10–24 heavy atoms, including well-known drug molecules, than those represented in its training data.The case studies specifically evaluated systems larger than the training molecules.
- Conclusions: 10-heavy-atom test molecules achieved RMSE versus DFT relative energies as low as 0.6 kcal/mol within 30 kcal/mol of each molecule’s minimum.The test set included 134 randomly selected molecules from GDB-11.
- Conclusions: ANI-1 was more accurate than DFTB and PM6 in the provided test cases when compared with the reference DFT level of theory.The comparison concerns the current ANI-1 version and the reported test cases.
- Conclusions: Single-point energies, and eventually forces, can be calculated as many as six orders of magnitude faster than DFT, with roughly force-field-equivalent per-atom scaling for very large molecules.The reported scaling is empirical.
- Conclusions: ANI-1 accuracy depends entirely on the training data, so adding molecules and atomic numbers is expected to improve accuracy and extend coverage to new chemical environments.The paper states that ANI could be built for other molecular classes and potentially crystals, while ANI-1 itself targets H, C, N, and O organic molecules.
Supplementary information for: “ANI-1: An extensible neural network
The supplementary information documents visualization examples, isomer structures, and datasets and parameters used to evaluate ANI-1. It also reports supplementary comparisons of ANI-1 with DFT and scaling behavior across training-set sizes.
- Supplementary figures: Figure S1 visualizes carbon atomic environment vectors for two formic-acid conformations that differ in their C-O-H angle.The vectors use modified angular symmetry functions with atomic-number differentiation.
- Supplementary figures: Figure S2 lists structural and geometric isomers used for the isomer case study and maps molecular indices to Figure 4’s isomer index.The mapping concerns the x-axis of Figure 4 in Section 4.2.
- Supplementary tables: Table S1 lists information and parameters used to generate the ANI-1 dataset, whose molecules are obtained from the GDB-11 database.Its first column records the number of heavy atoms per molecule in the test set, and “Total” combines all test sets.
- Supplementary tables: Table S2 compares ANI-1 and DFT absolute energies for 62 conformations of each of 134 randomly selected 10-heavy-atom molecules.The absolute-energy range is reported as -365,343 to -243,973 kcal/mol.
- Supplementary tables: Table S3 reports ANI-1 performance on 9171 normal-mode-sampling conformers from 134 randomly selected GDB-10 molecules.The imposed Ecap filter removes conformers above a molecule-specific energy threshold relative to the minimum.
- Supplementary tables: Table S4 examines ANAKIN-ME scaling with training-set size and compares two baseline methods trained on the same data.It reports training, validation, and full-test RMSE values alongside the percentage of 17.2 million data points used.