Source-linked AI summary
Atom-Density Representations for Machine Learning
Michael J. Willatt, Felix Musil, Michele Ceriotti
TL;DR
Atomic machine-learning models need representations that are both concise and sufficiently complete while respecting physical symmetries. The paper formulates smoothed atom-density environments in basis-independent bra-ket notation, recovering SOAP and related correlation-based representations. It establishes a unified framework in which tensor products, basis choices, and operators control structural, chemical, and target-property correlations.
Problem
Atomic representations must be complete and concise enough to support reliable property prediction while respecting translation, rotation, and permutation invariance.
Method
The paper represents atomic structures with smoothed atom densities and elemental kets, then uses symmetry averaging, tensor products, and linear operators to construct invariant feature representations and kernels.
Results
The framework unifies density-based representations, with radial-function and spherical-harmonic expansions giving a 1:1 mapping to SOAP kernel variants and real-space forms encoding body-order correlations.
Takeaways & Limitations
Different choices of tensor products, basis sets, and operators provide a systematic way to tune feature dimensionality and the representation of structural and chemical correlations.
Abstract
from arXiv · showhide
The applications of machine learning techniques to chemistry and materials science become more numerous by the day. The main challenge is to devise representations of atomic systems that are at the same time complete and concise, so as to reduce the number of reference calculations that are needed to predict the properties of different types of materials reliably. This has led to a proliferation of alternative ways to convert an atomic structure into an input for a machine-learning model. We introduce an abstract definition of chemical environments that is based on a smoothed atomic density, using a bra-ket notation to emphasize basis set independence and to highlight the connections with some popular choices of representations for describing atomic systems. The correlations between the spatial distribution of atoms and their chemical identities are computed as inner products between these feature kets, which can be given an explicit representation in terms of the expansion of the atom density on orthogonal basis functions, that is equivalent to the smooth overlap of atomic positions (SOAP) power spectrum, but also in real space, corresponding to $n$-body correlations of the atom density. This formalism lays the foundations for a more systematic tuning of the behavior of the representations, by introducing operators that represent the correlations between structure, composition, and the target properties. It provides a unifying picture of recent developments in the field and indicates a way forward towards more effective and computationally affordable machine-learning schemes for molecules and materials.
I. INTRODUCTION
The paper frames atomic-structure representations as a balance between concise, symmetry-compatible inputs and sufficient structural information for machine-learning predictions. It introduces a basis-independent, smoothed atom-density framework connecting established representations and enabling tunable feature construction.
- Machine-learning representations should preserve translation, rotation, and identical-atom permutation invariance because scalar properties are unchanged by these transformations.
- The formalism unifies density-based approaches by allowing different representations to arise from basis choices, atom-density expansions, and tunable couplings between feature components.The stated framework connects these choices to popular density-based representations and supports dimensionality reduction.
- The bra-ket formulation expresses kernels as basis-independent inner products between feature representations and connects them to kernel-ridge predictions.Reference-structure weights are determined by minimizing a loss against reference calculations.
- The framework represents structures as smooth atom-centered functions decorated with elemental kets, producing a unique but symmetry-variant description before averaging over symmetry groups.Smoothness supports smooth kernels and better-behaved regression.
- Direct translation averaging removes structural information, whereas averaging tensor products of the structure ket retains correlations while imposing symmetry invariance.The direct average retains only the number of atoms of each species.
C. Tensor-product representations
Tensor products determine how many-body information enters symmetry-invariant representations, while translational symmetrization naturally yields atom-centered environments and additive structure kernels. Linear operators can further tune positional and elemental correlations.
- C. Tensor-product representations: A linear kernel based on the translationally symmetrized ket can represent pair-potential-like behavior but cannot express properties requiring higher-order correlations.Taking the delta-function limit exposes the orientation-dependent pair-potential form.
- C. Tensor-product representations: Tensor products of structural or invariant kets introduce higher body-order correlations, with elementwise kernel powers providing an equivalent construction.A second-order construction makes learning depend simultaneously on two displacement vectors.
- D. Atom-centered representations: Grouping translationally invariant terms by an atom produces atom-centered contributions and makes the global kernel a sum or average of environment kernels.This additive relation follows directly from the density-based symmetrization.
- D. Atom-centered representations: A smooth cutoff restricts each atom-centered environment to a spherical neighborhood, providing localization that is useful for computational reasons.The cutoff can in principle be removed to include the entire structure.
- D. Atom-centered representations: Linear operators acting on position space, element space, or both can tune how the representation encodes structural and compositional relations.
E. Rotationally-invariant representations
Rotational invariance is obtained by averaging atom-centered environmental kets over the SO(3) rotation group. Because direct averaging loses angular correlations, tensor-product constructions provide richer invariant representations and support nonlinear kernels.
- A rotationally invariant representation is constructed by averaging the environmental ket over the SO(3) rotation group.
- Direct rotational averaging removes angular correlations around the environment center.
- Tensor products of environmental kets define a more general family of invariant representations that retains higher-order structural information.
- Nonlinear kernels correspond to tensor products of symmetrized kets and can be generalized using products built from different linear operators.
III. A UNIFIED PICTURE OF DENSITY-BASED REPRESENTATIONS
The formalism unifies density-based representations by encoding translation, rotation, and permutation symmetries in abstract invariant kets. Its real-space and basis expansions connect representation order to many-body correlations and to established structural fingerprints.
- Eq. (21) provides an abstract density-based representation that encodes translational, rotational, and permutation symmetries.
- Different basis choices yield alternative representations, including diffraction-related fingerprints for periodic structures.
- The ν = 1 invariant ket corresponds to spherical averaging, while ν = 2 captures three-body correlations involving two distances and an angle.
- Expanding the density shows that invariant kets represent (ν + 1)-body correlations among atoms within an environment.
- For sufficiently sharp atom densities, ν = 1 records pair distances, whereas ν = 2 adds pair distances and angles between atomic triplets.
- The invariant kets form a basis for representing arbitrarily complex invariant functions of atomic coordinates.
B. Behler-Parrinello symmetry functions
The formalism connects invariant density representations to pair distributions and Behler–Parrinello symmetry functions. These functions arise as projections of rotationally invariant kets onto suitable test functions.
- Pair distribution functions are directly connected to the corresponding density-based invariant representation.
- Behler–Parrinello symmetry functions can be viewed as projections of SO(3)-invariant kets onto suitable test functions.
- The two-body symmetry function G2(r) is obtained from the invariant representation through a weighted radial projection.
- Analogous expressions extend the construction to three-body symmetry functions G3(r, r′, ω).
C. Smooth Overlap of Atomic Positions
The formalism identifies SOAP as a basis representation of invariant kets, with the SOAP power spectrum corresponding to the ν = 2 feature vector. Raising normalized SOAP kernels to integer powers introduces tensor-product many-body structure.
- The SOAP power spectrum is an alternative representation of the invariant ket obtained by expanding environmental densities in radial functions and spherical harmonics.
- The ν = 2 feature vector corresponds to the SOAP power spectrum, while ν = 3 corresponds to the bispectrum.
- The bispectrum is used as a four-body feature vector in SOAP and SNAP for constructing accurate interatomic potentials through linear regression.
- The SOAP kernel is naturally an inner product between vectors representing a truncated power-spectrum expansion.
- Raising normalized SOAP kernels to an integer power ζ corresponds to tensor products of kets and introduces many-body character.
D. Tensorial Smooth Overlap of Atomic Positions (λ-SOAP)
The tensorial SOAP framework constructs rotationally covariant feature kets from smooth atomic densities and connects their inner products to λ-SOAP kernels and real-space correlations.
- Tensorial construction: Tensorial SOAP feature vectors use angular-momentum kets whose rotational behavior is controlled through spherical-harmonic expansions.Spherical harmonics enable explicit integration over rotations, while the angular ket is rotationally invariant only outside the atomic-environment subspace.
- Kernel connection: The inner product of these tensorial feature vectors reproduces the usual λ-SOAP kernel.The equivalence is stated for the inner product between the constructed vectors.
- Real-space representation: The real-space tensorial ket evaluates a three-body atom-density correlation and multiplies it by a spherical harmonic defined in the stencil reference frame.The stencil uses two density-evaluation points and an angular coordinate frame associated with (r, r′, ω).
- Higher-order correlations: Tensor products with λ = 0 kets increase explicit body-correlation order while preserving the desired rotational symmetry.This procedure has been found effective for increasing body-correlation order in tensorial SOAP.
E. Distributions vs sorted vectors
Sorted vectors and distributions provide closely related symmetry-aware representations: sorting removes permutation dependence, while the continuous sorted-vector representation is an inverse cumulative distribution function.
- Permutation invariance: Sorting permutation-variant structural descriptors makes the resulting vector permutationally invariant.The sorted vector can be related directly to the cumulative distribution function of the descriptor histogram.
- Distribution connection: The continuous representation of a sorted vector is the inverse cumulative distribution function associated with the histogram of its elements.This identifies sorted-vector and distribution-based descriptions as equivalent representations of the same element information.
- Distances: The Euclidean distance between sorted vectors is proportional to the L2 norm of the difference between their inverse cumulative distribution functions.Under the L1 norm, the corresponding distance is the one-dimensional earth mover’s distance.
- Information content: Incorporating physical symmetries produces representations that contain essentially the same information, although connections to sorted-vector descriptions can be less direct than connections among density-based representations.The relation may be complicated when permutation-variant quantities, such as overlap-matrix eigenvalues, are used.
- Feature transformation: Rotationally invariant representations can be transformed to improve regression quality or computational cost by incorporating chemical intuition and reducing feature dimensionality.The framework introduces rotationally invariant operators that couple geometric and elemental components in the SOAP power-spectrum basis.
- Feature transformation: A low-rank operator expansion can compress SO(3) fingerprints by retaining components selected from data correlations.Principal-component or sparse decompositions can identify linearly independent components and retain a chosen number according to covariance eigenvalues.
B. Radially-scaled kernels
Radially scaled and alchemical kernels modify SOAP features to emphasize chemically or spatially relevant correlations while reducing representation complexity.
- Radial scaling: Long-range regions can dominate environment overlaps despite weaker physical interactions, motivating low weights for distant atoms.The paper links this issue to the strong performance of multiscale kernels that assign very low weights to long-range components.
- Radial scaling: A rotationally invariant operator can implement radial scaling in the SOAP power spectrum and reproduce distance-weighted density representations.In a narrow-function, single-species approximation, the scaling is expressed through a radial weight and cutoff function.
- Radial scaling: For a Gaussian weighting function with zero central-atom weight, the construction becomes equivalent to two-body features from a prior descriptor.This connects the general density-based formalism to existing distance-scaled representations.
- Alchemical kernels: Alchemical projections couple chemical-element channels and can reduce the dimensionality of representations containing many species.The reduced-feature expression is more efficient and clarifies how elemental correlations and dimensionality reduction are introduced.
- Alchemical kernels: Off-diagonal chemical couplings can improve kernel-ridge property predictions, while reduced alchemical features make those couplings more efficient to compute.The operator formulation also allows projections to be optimized for a given regression problem.
- Alchemical kernels: For nsp species reduced to dJ ≪ nsp chemical channels, the feature dimensionality decreases by a factor of (nsp/dJ)^2.The full SOAP feature vector scales proportionally to nsp^2, whereas the reduced basis uses dJ channels.
D. Non-factorizable operators
Non-factorizable operators extend feature transformations beyond independent component-wise actions, while rotational invariance constrains their allowed couplings.
- General transformations: A further linear transformation can connect SOAP power-spectrum components to representations with more complicated internal-coordinate scaling functions.Because the transformation is linear, it also changes the effective regularization of ridge regression and can reduce feature dimensionality.
- Non-factorizable operators: Non-factorizable operators act jointly on tensor-product components rather than independently on each factor.For ν = 2, rotationally invariant operators are determined by their action on the corresponding basis vectors.
- Symmetry constraints: The operator must preserve rotational invariance, and any non-internal coordinate in the two-body operator must be cyclic.This constraint follows from requiring the transformed ket to retain the representation’s rotational symmetry.
- Coordinate scaling: Distance- and angle-based scaling can be represented with a diagonal operator and an adjustable scaling function.The paper identifies a three-body descriptor’s scaling function as an example and states that the construction extends to higher-body internal-coordinate functions.
- Feature selection: Linear contractions can select components whose contracted kernel remains close to the full kernel.Selected components may be determined by schemes such as CUR or farthest-point sampling.
E. Optimization of the density representation
The framework supports tuning density representations through operators that couple and scale feature channels, while dimensionality reduction can lower computational cost. General operator optimization risks overfitting, and systematic exploration and benchmarking remain future work.
- Optimization constraints: Optimizing the general Û operator involves many parameters, creating a concrete risk of overfitting.The risk is heightened because the optimized feature vector is subsequently used for regression.
- Optimization constraints: Dimensionality reduction can identify important linearly independent feature combinations before optimizing Û against target properties.This ordering is intended to make subsequent target-property optimization less likely to overfit.
- Open directions: Systematic exploration of representation choices and benchmarking across regression problems are left for future work.The proposed scaling coefficients are intended to make representations better suited to regression for a target property.
- Representation design: Different tensor-product and symmetrization choices produce representations that capture varying amounts of inter-atomic correlation.The atom-density formulation is basis-independent and can therefore express multiple correlation orders within one framework.
- Representation design: The framework unifies density-based representations by showing that popular methods arise from different limits or basis choices for invariant kets.With radial functions and spherical harmonics, symmetrized kets map one-to-one onto SOAP-kernel variants.
- Representation design: Operators can couple and scale different representation channels, reducing dimensionality while modifying the SOAP power spectrum representation.The framework presents these modifications as general ways to optimize density-based representations.