Source-linked AI summary
Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models
Luis Barroso-Luque, Muhammed Shuaibi, Xiang Fu, Brandon M. Wood, Misko Dzamba, Meng Gao, Ammar Rizvi, C. Lawrence Zitnick, Zachary W. Ulissi
TL;DR
Materials ML research lacks sufficiently open, large, and diverse training data and pretrained models for efficiently exploring the enormous materials space. The paper releases OMat24 and pretrained models, showing strong benchmark performance and reduced systematic softening across architectures. The resource supports further open development of AI-assisted materials science, while remaining limited by its DFT approximations, bulk-only coverage, and magnetic-ordering assumptions.
Problem
Materials discovery spans an enormous search space, while progress in training generalizable MLIPs is constrained by limited access to large, diverse, and reproducible training data.
Method
The paper releases OMat24, a large open dataset of diverse non-equilibrium inorganic-material configurations, together with pretrained EquiformerV2 and eSEN models.
Results
OMat24 pretraining yields substantially better Matbench Discovery results across nearly all metrics and reduces or eliminates systematic softening across model architectures.
Takeaways & Limitations
The open dataset and models provide a foundation for reproducible MLIP research and further advances in AI-assisted materials science.
Takeaways & Limitations
OMat24 uses PBE and PBE+U calculations, covers only periodic bulk structures, and generally represents magnetic materials through specific initialized magnetic configurations rather than guaranteed ground states.
Abstract
from arXiv · showhide
The ability to discover new materials with desirable properties is critical for numerous applications from helping mitigate climate change to advances in next generation computing hardware. AI has the potential to accelerate materials discovery and design by more effectively exploring the chemical space compared to other computational methods or by trial-and-error. While substantial progress has been made on AI for materials data, benchmarks, and models, a barrier that has emerged is the lack of publicly available training data and open pre-trained models. To address this, we present a Meta FAIR release of the Open Materials 2024 (OMat24) large-scale open dataset and an accompanying set of pre-trained models. OMat24 contains over 110 million density functional theory (DFT) calculations focused on structural and compositional diversity. Our EquiformerV2 models achieve state-of-the-art performance on the Matbench Discovery leaderboard and are capable of predicting ground-state stability and formation energies to an F1 score above 0.9 and an accuracy of 20 meV/atom, respectively. We explore the impact of model size, auxiliary denoising objectives, and fine-tuning on performance across a range of datasets including OMat24, MPtraj, and Alexandria. The open release of the OMat24 dataset and models enables the research community to build upon our efforts and drive further advancements in AI-assisted materials science.
1 Introduction
Materials discovery requires efficient screening across an enormous chemical space, while ML model development is constrained by limited access to large, diverse training data. OMat24 addresses this gap with an open, large-scale dataset and pretrained models that improve generalization and reduce systematic softening across architectures.
- 1 Introduction: An enormous materials search space makes efficient screening essential for applications including energy storage, carbon-neutral fuels, and direct CO2 capture.ML models may improve computational screening relative to traditional DFT by reducing the cost of materials simulations.
- 1 Introduction: Open training data is needed because inaccessible underlying datasets prevent independent replication and broader open-science use of materials models.
- 1 Introduction: OMat24 provides an open dataset and pretrained models designed to improve MLIP generalization across a wide range of inorganic materials.The dataset samples diverse non-equilibrium configurations and elemental compositions, while pretrained EquiformerV2 and eSEN models are fine-tuned for differing DFT settings.
- 1 Introduction: OMat24-trained models reduce or eliminate systematic softening, with improvements consistent across different model architectures.Systematic softening refers to underprediction effects affecting energies, forces, and phonons.
2 Results
OMat24 is a large, diverse, openly released DFT dataset designed to improve MLIP generalization across equilibrium and non-equilibrium inorganic materials. Models trained on OMat24 achieve leading Matbench-Discovery performance and reduce systematic softening, while dataset scope and DFT-setting differences remain important boundaries.
- 2 Results: OMat24 is up to 2 orders of magnitude larger than other openly available datasets, spans diverse compositions and far-from-equilibrium structures, and is released under a permissive CC-by license.The dataset contains broad periodic-table coverage and is intended for community reuse.
- 2.1.1 OMat24 train, validation, and test splits: The WBM test split produces the highest prediction errors, indicating the strongest distribution shift and a more informative test of model generalization.The split is separated by unique anonymized formulas and space groups.
- 2.1.2 Dataset limitations: OMat24 is limited to periodic bulk structures and PBE/PBE+U calculations, excluding point defects, surfaces, non-stoichiometry, lower-dimensional structures, and reliable magnetic ground-state sampling.Formation-energy comparisons also show that most MP–OMat24 differences are under 20 meV/atom, but deviations above 50 meV/atom mainly arise for magnetic compounds with different initialized orderings.
- 2.2.1 Model performance on Matbench-Discovery: Pre-training on OMat24 yields substantially better Matbench-Discovery results across nearly all metrics, with eSEN achieving an F1 score of 0.925 and an energy-above-hull MAE of 18 meV/atom.These models set the reported best performance at the time of writing after fine-tuning for DFT-setting differences.
- 2.2.1 Model performance on Matbench-Discovery: The five top-performing Matbench-Discovery models trained using OMat24 achieve energy-above-hull errors near 20 meV/atom, while energy-conserving models show the lowest thermal-conductivity errors.Direct-force models perform poorly on phonon prediction tasks, as reflected by higher thermal-conductivity SRME.
- 2.2.2 Model evaluation of systematic softening: OMat24 training reduces or eliminates systematic underprediction of energies, forces, and phonons across architectures, although fine-tuning with relaxation data can reintroduce some softening.OMat24-only models show less softening across all orders, while OAM models substantially reduce zeroth- and first-order softening relative to MPtrj-only models.
3 Discussion
OMat24 provides a large, diverse PBE-DFT foundation for ML interatomic potentials while retaining important scope and fidelity limitations. The discussion points toward higher-fidelity functionals and information-efficient sampling as next steps.
- OMat24’s broad PBE-DFT coverage supports MLIP development, but excludes transition-state frames and highly off-equilibrium regions such as bond breaking or large coordination changes.The dataset nevertheless supports reported extrapolation to defects, grain boundaries, adsorption, and cleavage energies.
- PBE and PBE+U introduce approximation errors, while bulk, defect-free structures and magnetic-ordering exclusions constrain prediction reliability and scope.These limitations particularly affect materials with strong on-site Coulomb interactions in d and f orbitals.
- Higher-fidelity datasets using SCAN, r2SCAN, or other functionals, combined with physically motivated sampling, are identified as important next steps.Suggested sampling directions include active learning, molecular dynamics, and metadynamics beyond local minima.
- OMat24 can support multi-fidelity, delta-learning, and fine-tuning approaches that combine extensive low-level force data with smaller high-level datasets.The stated goal is reduced computational cost for higher-accuracy MLIPs.
- OMat24 also provides a platform for sampling and compression studies aimed at identifying non-redundant physical information and training accurate models with less data.This direction is framed as a way to investigate data efficiency rather than as a demonstrated result of the dataset itself.
- The authors characterize OMat24 as a resource for highly accurate MLIPs with reduced systematic softening biases and continued community development.
4 Methods
OMat24 is constructed from diverse non-equilibrium structures generated from Alexandria relaxations and labeled with DFT calculations. The study trains EquiformerV2 and eSEN models under multiple dataset, denoising, and fine-tuning strategies while evaluating softening and related properties.
- 4.1 OMat24 structure generation and sampling: OMat24 samples non-equilibrium structures using Boltzmann-rattled structures, AIMD, and rattled relaxations initialized from randomly selected Alexandria bulk structures.The starting-point choice preserves compositional diversity while avoiding configurations too far from equilibrium for DFT convergence.
- 4.1 OMat24 structure generation and sampling: The sampling strategies increase structural diversity, but the authors note that many alternative strategies remain possible and active-learning comparisons at scale are unresolved.
- 4.2 OMat24 DFT calculation settings and details: DFT calculations use VASP with periodic boundary conditions, PAW pseudopotentials, PBE exchange-correlation, and Hubbard U corrections for specified oxide and fluoride materials.
- 4.2 OMat24 DFT calculation settings and details: AIMD runs use 50 steps with a 2 fs timestep, which the authors identify as large for typical hydrogen-containing simulations because the goal is diverse sampling rather than exact trajectory integration.
- 4.3 OMat24 MLIP training strategies: The models predict energy, forces, and stress using EquiformerV2 and energy-conserving eSEN architectures, with experiments spanning OMat24-only, MPtrj-only, and fine-tuned strategies.Denoising objectives are varied, and fine-tuning uses MPtrj or leakage-controlled sAlexandria data.
- 4.4 Softening evaluation: Softening is evaluated through predicted energies and forces on WBM high-energy states and phonon frequencies on approximately 10,000 PBE MDR structures.The evaluation uses architecture-specific ASE calculators and a relaxation-based phonon protocol.
Code Availability
The paper releases training and architecture code together with pretrained EquiformerV2 and eSEN checkpoints for community use.
- Training and architecture code is available on FAIRChem GitHub, while pretrained EquiformerV2 and eSEN checkpoints are downloadable from Hugging Face.The checkpoints are compatible with fairchem-core 1.10.
Supplementary Information
The supplementary information documents differences between OMat24 and Materials Project pseudopotential choices. Table A.1 lists the affected element symbols and generation dates.
- OMat24 uses VASP pseudopotential version 54, whereas Materials Project calculations use older releases, with additional symbol differences for Yb and W.
- Table A.1 records the element symbols and pseudopotential generation dates that differ between OMat24 and Materials Project PBE calculations.
B Dataset statistics
OMat24 combines broad structural, compositional, and local-environment diversity with varied energy, force, and stress distributions. Supplementary analyses also document AIMD drift, convergence behavior, and the corrections needed to compare OMat24 energies with experiment and other DFT settings.
- Dataset distributions: OMat24 has broader label distributions than relaxation-only datasets, including higher energies and wider force and stress ranges than MPtrj.The dataset was designed to represent both equilibrium and non-equilibrium properties.
- AIMD characteristics: AIMD subsets show narrower force and stress distributions and lower energies, while reported AIMD drift forces can reach up to 1 meV/Å under OMat24 settings.The dataset analysis compares AIMD with rattled-relaxation and rattled-Boltzmann subsets and examines drift using the OMat24-1M test set.
- Elemental and structural diversity: OMat24 and Alexandria have more uniform elemental coverage than MPtrj, while OMat24 contains greater local-environment diversity through higher, more uniform element-pair counts.Element-pair comparisons use neighborhoods within 3.5 Å.
- DFT convergence: Models trained on OMat24 predict tightly converged DFT forces with test errors effectively equivalent to those on regularly converged calculations.This comparison is reported for eSEN and EquiformerV2-S on both convergence settings.
- Formation-energy corrections: The OMat24 compatibility-correction fit yields a 55 meV/atom mean absolute error against experimental formation energies for 222 compounds.The corrections follow the Materials Project-style GGA/GGA+U mixing and anion-correction procedure.
- Formation-energy corrections: Formation energies from OMat24-trained models should use OMat24 elemental references and compatibility corrections appropriate to the DFT setting.The authors emphasize that MP2020-style corrections require attention to the relevant correction scheme and calculation context.
D Dataset Learning Curves
eSEN learning curves show decreasing energy and force errors as OMat24 training size grows, with performance plateauing near the full dataset scale. The learning-curve visual tracks both errors against the number of training structures.
- Learning-curve result: Prediction errors decrease with more training structures and plateau at approximately 100 million samples for eSEN.The plateau matches the total size of the OMat24 training set in these experiments.
E Systematic softening improvements across model architectures
OMat24-related improvements in systematic softening extend across additional architectures, supporting an architecture-independent association with dataset diversity. The analysis covers energy, force, and phonon softening distributions.
- Cross-architecture result: Softening improvements appear consistently across five architectures, including GRACE and Sevenn, across energy, force, and phonon orders.The authors attribute the consistency primarily to OMat24 dataset diversity rather than a single architecture.
F Effects of denoising augmentation versus dataset diversity
Denoising augmentation improves MPtraj-trained models, but OMat24’s larger in-domain diversity produces the strongest gains and broader accuracy away from the convex hull. With OMat24, adding DeNS instead incurs a slight validation penalty.
- Denoising augmentation: DeNS improves eSEN and EquiformerV2 models relative to non-DeNS counterparts and other third-party models on the Matbench Discovery task.For larger EquiformerV2 models, DeNS also supports effective training on the smaller MPtraj dataset.
- Model size: EquiformerV2 energy MAE remains 35–36 meV/atom across model sizes, indicating practical utility for the smaller model.This result is reported for compliant models in the Matbench Discovery comparison.
- Denoising augmentation: MPtraj models show an overall improvement when trained with DeNS rather than MPtraj alone.The comparison is presented for energy-above-hull prediction error using a 40 meV/atom rolling window.
- Dataset diversity: OMat24 pretraining produces the most significant performance improvements and broader constant-accuracy windows away from the convex hull.The broader window indicates improved extrapolation to lower-density regions of the benchmark distribution.
- Denoising augmentation: Adding DeNS to OMat24 training causes a slight validation-performance penalty for EquiformerV2-S models.The authors interpret this as indicating that OMat24’s dataset diversity does not benefit from the same augmentation or regularization.
H.1 OMat24 test metrics
This section reports OMat24 test metrics for EquiformerV2 models trained on the OMat24 dataset, including energy, force, and stress errors.
- Table H.1 lists OMat24 test mean absolute error metrics for EquiformerV2 models trained on the OMat24 dataset.
- The reported errors cover energy in meV/atom, forces in meV/Å, and stress in meV/Å3.
I MPtrj validation metrics
This section reports MPtrj validation metrics for models trained solely on MPtrj and for OMat24-pretrained models fine-tuned on MPtrj.
- Tables I.1 and I.2 report validation metrics for models trained on MPtrj and for OMat24-pretrained models fine-tuned on MPtrj.
- The MPtrj validation metrics use energy, force, and stress mean absolute errors measured in meV/atom, meV/Å, and meV/Å3, respectively.
- Table I.2 specifically reports validation metrics for fine-tuning OMat24-pretrained models on MPtrj.