Source-linked AI summary

Error estimates for solid-state density-functional theory predictions: an overview by means of the ground-state elemental crystals

Kurt Lejaeghere, Veronique Van Speybroeck, Guido Van Oost, Stefaan Cottenier

arXiv:1204.2733v5cond-mat.mtrl-sciphysics.comp-phphysics.data-an

TL;DR

DFT is increasingly used to predict material properties, but reliable expected-error estimates are scattered and uncommon. The review synthesizes benchmark evidence from elemental crystals into a practical protocol and assesses reproducibility across implementations. It finds that intercode differences are much smaller than typical theory–experiment differences, supporting common PBE error estimates across the compared approaches.

  • Problem

    Reliable expected-error information for DFT predictions is scattered across the literature, although such predictions increasingly inform experiments.

  • Method

    The review statistically compares DFT predictions with experiment for ground-state elemental crystals and evaluates method and code dependence using a quality factor Δ.

  • Results

    The quality factor is Δ = 1.9 meV/atom for PAW(VASP) versus APW+lo(WIEN2k) and Δ = 3.3 meV/atom for PAW(GPAW) versus APW+lo(WIEN2k), about an order of magnitude below the theory–experiment gap.

  • Takeaways & Limitations

    The reported intrinsic systematic deviations and residual error bars can be applied to PBE predictions regardless of the compared computational approach.

  • Takeaways & Limitations

    The review is a starting point: its statistical procedure should be extended to other functionals and methods, while experimental error bars were beyond its scope.

Abstract

from arXiv · show

Predictions of observable properties by density-functional theory calculations (DFT) are used increasingly often in experimental condensed-matter physics and materials engineering as data. These predictions are used to analyze recent measurements, or to plan future experiments. Increasingly more experimental scientists in these fields therefore face the natural question: what is the expected error for such an ab initio prediction? Information and experience about this question is scattered over two decades of literature. The present review aims to summarize and quantify this implicit knowledge. This leads to a practical protocol that allows any scientist - experimental or theoretical - to determine justifiable error estimates for many basic property predictions, without having to perform additional DFT calculations. A central role is played by a large and diverse test set of crystalline solids, containing all ground-state elemental crystals (except most lanthanides). For several properties of each crystal, the difference between DFT results and experimental values is assessed. We discuss trends in these deviations and review explanations suggested in the literature. A prerequisite for such an error analysis is that different implementations of the same first-principles formalism provide the same predictions. Therefore, the reproducibility of predictions across several mainstream methods and codes is discussed too. A quality factor Delta expresses the spread in predictions from two distinct DFT implementations by a single number. To compare the PAW method to the highly accurate APW+lo approach, a code assessment of VASP and GPAW with respect to WIEN2k yields Delta values of 1.9 and 3.3 meV/atom, respectively. These differences are an order of magnitude smaller than the typical difference with experiment, and therefore predictions by APW+lo and PAW are for practical purposes identical.

I. INTRODUCTION

DFT is widely used to interpret measurements and predict material properties, making quantitative error estimates increasingly important. This review addresses intrinsic and numerical errors through benchmark comparisons and standardized property calculations.

  • I. INTRODUCTION: DFT predictions increasingly guide interpretation of measurements and planning of future experiments, but quantitative error estimates remain uncommon.When theory and experiment disagree, users often resort to higher-order theories instead of estimating expected DFT deviations.
  • I. INTRODUCTION: The review quantifies intrinsic and numerical DFT errors and examines physical explanations for the observed deviations.Benchmark studies compare implementations and functionals across materials and properties.
  • I. INTRODUCTION: Ground-state elemental crystals provide a broad, interpretable test set with accessible and relatively reliable experimental data.Periodic-table classification supports visualization and interpretation of the results.
  • I. INTRODUCTION: The calculations cover cohesive energy, equilibrium volume, bulk modulus, its pressure derivative, and elastic constants derived from total-energy or stress calculations.Equilibrium quantities can be obtained by optimization or equation-of-state fitting; elastic constants require stress tensors and deformed geometries.
  • I. INTRODUCTION: Bulk modulus and its pressure derivative are numerically sensitive because they depend on curvature and higher derivatives of the equation of state.B1 is obtained by fitting calculated E(V) data and is more sensitive than the bulk modulus.

B. Comparing theory and experiment

The comparison requires matching theoretical and experimental conditions, including corrections for zero-point and thermal effects. The review uses established correction schemes while retaining certain quantities uncorrected when data or reliable estimates are limited.

  • Experimental values are extrapolated to 0 K and corrected for zero-point vibrations before comparison with standard DFT predictions.For cohesive energies, the zero-point contribution is added to the experimental value under the chosen sign convention.
  • Equilibrium volumes receive thermal-expansion and zero-point volume-shift corrections, using experimental expansion coefficients when available.When expansion data are unavailable, the coefficient can be estimated from the moleculization energy, with DFT values used if both experimental inputs are missing.
  • B0 is sensitive to thermal expansion and zero-point vibrations, so both effects are considered through separate corrections.The bulk modulus is linked to the curvature of E(V), making narrow volume ranges and shallow equations of state numerically problematic.
  • B2 is highly sensitive, difficult to extract from a few E(V) points, and absent from the third-order Birch-Murnaghan equation used in the main fit.The study therefore uses the intrinsic Birch-Murnaghan value rather than applying a higher-order fit.
  • B1 is not adjusted for thermal expansion or zero-point effects because these modifications are often negligible relative to experimental or computational errors.The review notes that pressure derivatives such as B1 are already difficult to measure or determine from first principles.
  • Thermal corrections are not applied to the elastic constants Cij because experimental data for their pressure derivatives are scarce.The review notes that analogous modifications could nevertheless be envisioned for these constants.

A. Test set preparation

The benchmark set is built from ground-state elemental crystal structures, with selected low-temperature alternatives and individually optimized geometries. These structures are then used to compare PBE predictions with experiment across the reviewed properties.

  • A. Test set preparation: The test set uses ground-state elemental crystals at 0 K, replacing literature structures for boron, nitrogen, oxygen, and sulfur when lower-temperature phases are suggested.The selected structures are summarized in Table I.
  • A. Test set preparation: 13 crystal volumes are calculated and fitted to a third-order Birch-Murnaghan equation of state, with expanded ranges for shallow E(V) curves.The atomic positions and cell shape are optimized separately at each volume, followed by reoptimization at the fitted equilibrium volume.
  • A. Test set preparation: The optimized structures form the definitive test set and are submitted to the COD and ICSD crystallographic databases.
  • A. Test set preparation: Most reviewed properties are determined for each structure to quantify differences between PBE and experimental values.The DFT comparison is performed with VASP using settings that converge energy differences to at most a few meV per atom.
  • A. Test set preparation: The elastic-constant comparison uses previously published data for most of the benchmark set, with alternative experimental values selected for Ba.The cited prior study considered only bcc, fcc, and hcp structures, affecting coverage for Li and Na.

1. Linear regression

The review replaces simple absolute-error summaries with a regression-based separation of systematic deviation and residual scatter. This approach treats relative shifts as more appropriate for strictly positive properties and models experiment as a scaled theory value plus random error.

  • 1. Linear regression: For strictly positive quantities, the analysis treats relative deviations rather than assuming a constant absolute offset across scales.
  • 1. Linear regression: Figure 1 compares the regression line X = βT with the bisector X = T to visualize systematic differences.Green arrows represent the residual error bar, defined as the standard deviation of regression errors (SER).
  • 1. Linear regression: Linear regression separates systematic deviations from residual scatter when comparing DFT predictions with experiment.The slope captures the systematic trend, while the scatter around the regression line quantifies remaining error.
  • 1. Linear regression: The regression model is X = βT + ϵ, with perfect theory–experiment correlation distorted by zero-mean random error.
  • 1. Linear regression: DFT scatter becomes apparent across many compounds because some crystal subsets are described better by the functional than others.Similar systems within a material family can behave almost identically, distinguishing intrinsic DFT error bars from experimental error bars.

2. Eliminating outliers

The review removes entire element subsets that systematically distort DFT–experiment agreement, grouping outliers by shared physical properties rather than deleting individual materials.

  • 2. Eliminating outliers: The periodic-table classification enables visual interpretation of deviations and organizes materials with related physical behavior.The review uses the conventional medium-form periodic table to preserve intuitive recognition.
  • 2. Eliminating outliers: Outlier removal targets whole structure types, preventing well-behaved materials from being retained for potentially misleading reasons.The procedure excludes groups associated with poorly described physical mechanisms, avoiding bias toward smaller errors.
  • 2. Eliminating outliers: Eight subsets classify elemental crystals by common physical properties, including bonding, magnetism, correlation, coordination, dispersion, and crystal character.The categories include alkali and alkaline earth metals, transition metals, magnetic and correlation-dominated materials, p-block compounds, molecular crystals, and noble gases.
  • 2. Eliminating outliers: Subsets are iteratively excluded when at least half their elements significantly depart from the dominant trend, using a two-sided p-value threshold of 10%.After each elimination, the significance criterion is recalculated until no deviating subsets remain.

3. Predicting experiment

The review converts DFT–experiment discrepancies into regression-based systematic corrections and residual error bars, while noting that error behavior depends on property scale and material type.

  • 3. Predicting experiment: Relative systematic deviations and residual error bars summarize intrinsic errors after least-squares regression against experiment.The regression value is obtained by correcting the theoretical result using xreg = xth −(1−β)xth.
  • 3. Predicting experiment: p = 6 · 10−11 for equilibrium volume and p = 5 · 10−4 for bulk modulus show significant deviations from xexp = xth.These results support incorporating systematic deviations rather than relying only on residual scatter.
  • 3. Predicting experiment: Bond type influences residual errors: strong covalent and half-filled d-shell bonds are described better and tend to produce more compact structures.Other crystal subsets with similar volumes can still show larger relative errors.
  • 3. Predicting experiment: 438±15 GPa is obtained for diamond’s 0 K bulk modulus after systematic, zero-point, and intrinsic-error corrections to a VASP-PBE result.The uncorrected calculation is 434.8 GPa; the systematic PBE correction changes it to 456.0 GPa before a −17.6 GPa zero-point correction.
  • 3. Predicting experiment: 0.37 σV (0.40 ˚A3/atom) is the mean absolute difference after regression correction for GaAs, compared with 0.72 σV (0.78 ˚A3/atom) without it.The comparison uses 31 theoretical predictions and illustrates the practical value of applying systematic deviations.

1. Errors per materials type

PBE errors vary by material class because some physical phenomena are poorly described, while excluding identified outliers yields more representative error estimates. Dispersion-governed and magnetic materials are prominent problematic subsets.

  • Dispersion-governed materials: Dispersion-governed compounds are poorly described by regular GGA functionals because they omit London forces, reducing cohesion and inflating volumes.Affected crystals include noble gases, dimeric crystals, graphite, and sulfur; the importance of dispersion depends on both element type and crystal structure.
  • Magnetic materials: Magnetic compounds remain difficult for current GGA functionals, with magnetic-state approximations and thermal extrapolation effects contributing to deviations from experiment.Manganese and chromium illustrate these issues through intricate magnetic states and phase-transition effects near Curie or Néel temperatures.
  • Outlier treatment: PBE error estimates depend strongly on the property being considered after affected materials are excluded from the test set.The review uses these restricted subsets to characterize residual errors and systematic deviations for different properties.
  • Experimental accuracy: Experimental accuracy also affects correspondence for individual compounds, especially for higher-order properties measured less precisely than lattice constants or cohesive energies.The review notes that elastic constants and their derivatives generally have lower experimental precision.
  • Volumes and elastic properties: Equilibrium volumes show an almost perfect correlation after outlier removal but are consistently approximately 4% too large because GGA underestimates bond strength.Bulk moduli are also underestimated, with intrinsic residual errors usually below 10 to 15%, because weaker binding makes structures more compressible.
  • Volumes and elastic properties: The bulk modulus derivative B1 is overestimated and remains less strongly correlated with experiment after outlier removal, with correlation coefficient r = 0.849.The review attributes this behavior to GGA underbinding, which changes the curvature of the equation of state more rapidly.

IV. NUMERICAL ERRORS

The numerical-error analysis tests whether independent DFT implementations reproduce the intrinsic-error benchmarks. It uses a common PBE setup and ground-state elemental crystals to compare codes systematically.

  • Purpose and benchmark: The numerical-error section examines whether different DFT implementations provide reproducible predictions, a prerequisite for using intrinsic-error estimates.The benchmark uses ground-state elemental crystals to assess how implementation choices affect calculations.
  • Reference implementation: WIEN2k, an all-electron APW+lo code, serves as the numerical-accuracy reference for a given functional when sufficiently large basis sets and dense k-meshes are used.The assessment compares other codes against this reference implementation.
  • Common computational protocol: All three codes use the same PBE functional, and properties are extracted from a 7-point equation of state spanning 0.94 V0 to 1.06 V0.The procedure starts from the VASP-optimized ground-state crystal and extracts E0, V0, B0, and B1 while keeping geometries frozen.
  • Common computational protocol: The calculations use scalar-relativistic treatments without spin-orbit contributions to avoid an additional algorithm implementation and apply the procedure uniformly across elements.This choice makes the computational setup more uniform for the code comparison.
  • Test-set modifications: Mn and S are assigned simplified but physically relevant crystal structures to reduce computational effort while retaining representative benchmark cases.Mn is treated in an antiferromagnetic fcc phase, while sulfur uses the β Po phase.
  • Test-set modifications: All other elements retain their previously VASP-optimized structures, preserving the diversity of the elemental code-benchmark set.The corresponding crystal CIF files are available in the Supplementary Material.

B. Agreement between implementations

The review compares implementations through their energy-versus-volume curves rather than isolated equation-of-state parameters. The resulting Δ factor provides a single energy-distance measure for intercode reproducibility.

  • Definition of Δ: A single numerical-error value is obtained by a weighted average because the benchmark properties have different units and characteristic relative-error scales.The scale is mainly determined by the computational procedure and remains comparable between code–experiment and intercode assessments.
  • Definition of Δ: The Δ factor compares E(V) curves directly, avoiding inflated relative errors when shallow curves produce large parameter differences despite similar equations of state.For dispersion-governed compounds, the curves may differ by no more than a few meV per atom even when fitted parameters differ substantially.
  • Interpretation and normalization: Δ is interpreted as the root-mean-square energy difference between two programs’ E(V) curves, averaged over all elemental crystals.All equations of state are shifted to zero at their equilibrium volume because codes may use different reference energies.
  • Reported quantities: The benchmark table reports V0 in Å3/atom, B0 in GPa, and Δi in meV/atom, while B1 is dimensionless.These units define the presentation of code comparisons for representative cases.
  • Results: The PAW comparisons give Δ = 3.3 meV/atom for GPAW and an experimental-reference difference of 23.5 meV/atom for PAW(VASP).The intercode spread is therefore much smaller than the DFT–experiment difference, although GPAW and VASP are not entirely identical.
  • Practical implication: The Δ-factor framework can also guide users in selecting methods for tasks where the accuracy of energy-versus-volume relations matters.The review describes this as a practical use of reproducibility information alongside insight into intrinsic-error reproducibility.

V. CONCLUSIONS

The review quantifies intrinsic and numerical errors in DFT-PBE predictions using ground-state elemental crystals and develops practical error estimates for solid-state properties. Numerical differences between codes are much smaller than typical theory–experiment deviations, supporting transfer of intrinsic error estimates across implementations while marking clear scope limits.

  • DFT-PBE errors were reviewed on ground-state elemental crystals for five energetic and elastic materials properties.The properties include cohesive energy and elastic quantities V0, B0, B1, and Cij.
  • Systematic deviations and residual error bars were quantified by statistically comparing DFT predictions with experiment.The analysis reproduces known GGA traits, including typical underbinding, and supports a recipe for correcting PBE results and attaching error estimates.
  • The benchmark’s conclusions should be extended with binary ionic compounds and transition-metal oxides, and the current review does not incorporate experimental error bars.The authors describe the accuracy review as a starting point and propose broader benchmark sets and further statistical extensions.
  • 1.9 meV/atom and 3.3 meV/atom are the quality factors Δ for PAW(VASP) and PAW(GPAW), respectively, relative to APW+lo(WIEN2k).These quality factors express average rms energy differences between equations of state from the compared codes.
  • Numerical errors are an order of magnitude smaller than the gap between DFT predictions and experiment, allowing intrinsic PBE estimates to apply across computational approaches.This transferability is useful when interpreting DFT results in an experimental context.
Loading 1204.2733v5…