Source-linked AI summary

The Open Catalyst 2020 (OC20) Dataset and Community Challenges

Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, Muhammed Shuaibi, Morgane Riviere, Kevin Tran, Javier Heras-Domingo, Caleb Ho, Weihua Hu, Aini Palizhati, Anuroop Sriram, Brandon Wood, Junwoong Yoon, Devi Parikh, C. Lawrence Zitnick, Zachary Ulissi

arXiv:2010.09990v5cond-mat.mtrl-scics.LG

TL;DR

Catalyst modeling needs efficient methods that generalize across diverse compositions and chemistries, but catalysis datasets remain relatively small. The paper introduces the OC20 dataset, tasks, and graph-neural-network baselines, finding that hybrid relaxation approaches outperform direct prediction while baseline accuracies remain below practical usefulness.

  • Problem

    Heterogeneous catalyst modeling spans complex compositions and chemistries, while catalysis datasets remain smaller than those in related fields, limiting available training data.

  • Method

    The paper constructs the OC20 DFT dataset with three catalyst-modeling tasks, predefined generalization splits, and graph-neural-network baseline models.

  • Results

    Hybrid relaxation approaches outperformed direct approaches across all metrics, but predicted energies within a tight threshold reached only 2%–6%, below practical usefulness.

  • Takeaways & Limitations

    OC20 provides an open benchmark and baseline framework for comparing models across catalyst properties and generalization settings.

  • Takeaways & Limitations

    Only 18.9% of sampled compositions succeeded, and the calculations covered approximately 0.07% of possible sites, leaving the dataset severely sparse.

Abstract

from arXiv · show

Catalyst discovery and optimization is key to solving many societal and energy challenges including solar fuels synthesis, long-term energy storage, and renewable fertilizer production. Despite considerable effort by the catalysis community to apply machine learning models to the computational catalyst discovery process, it remains an open challenge to build models that can generalize across both elemental compositions of surfaces and adsorbate identity/configurations, perhaps because datasets have been smaller in catalysis than related fields. To address this we developed the OC20 dataset, consisting of 1,281,040 Density Functional Theory (DFT) relaxations (~264,890,000 single point evaluations) across a wide swath of materials, surfaces, and adsorbates (nitrogen, carbon, and oxygen chemistries). We supplemented this dataset with randomly perturbed structures, short timescale molecular dynamics, and electronic structure analyses. The dataset comprises three central tasks indicative of day-to-day catalyst modeling and comes with pre-defined train/validation/test splits to facilitate direct comparisons with future model development efforts. We applied three state-of-the-art graph neural network models (CGCNN, SchNet, Dimenet++) to each of these tasks as baseline demonstrations for the community to build on. In almost every task, no upper limit on model size was identified, suggesting that even larger models are likely to improve on initial results. The dataset and baseline models are both provided as open resources, as well as a public leader board to encourage community contributions to solve these important tasks.

Introduction

The introduction motivates OC20 as a response to the computational cost, chemical complexity, limited dataset scale, and weak generalization of machine-learning approaches for heterogeneous catalysis. It presents OC20 as a large, diverse, publicly available dataset with accompanying challenges for catalyst-model development.

  • Motivation: Catalysis supports renewable electricity, fuel production, and ammonia synthesis, but catalyst discovery and optimization remain time-intensive.These applications include fuel cells, renewable-resource fuels, and fertilizer production.
  • Motivation: Machine-learning models trained on expensive DFT data offer a potentially efficient approach to molecular simulation, but heterogeneous catalysis spans complex organic and inorganic chemical spaces.The space includes diverse molecules, conformers, elements, coordination environments, lattice structures, and long-range interactions.
  • Data limitations: Catalysis datasets have expanded from O(100) to O(100,000) relaxations, yet models continue improving with more data and show limited generalization.The introduction attributes the smaller dataset sizes to catalysis’s additional complexity and computational cost.
  • OC20 dataset: Over 1.2 million DFT relaxations and ca. 250 million single-point calculations comprise OC20 across a substantially larger structure and chemistry space than previously realized.OC20 is intended as a crucial stepping stone for machine-learning models targeting practical catalysis applications.
  • OC20 dataset and challenges: OC20 includes 82 adsorbates and surfaces from 55 elements and mixtures, enriches calculations with electronic-structure and dynamics data, and supports three open prediction challenges with train/validation/test splits.The challenges address energy and force prediction, nearby relaxed-state prediction, and relaxed adsorption-energy prediction.

Tasks

The dataset defines three related catalyst-modeling tasks around DFT-based structure relaxation: predicting energies and forces, relaxed structures, or relaxed energies from initial structures. All tasks use periodic surface–adsorbate structures with DFT-computed ground truth.

  • Task relationships: The tasks are related because success in one may aid the others, while the dataset remains open to additional future tasks and community input.The authors describe these as three related tasks rather than the only possible tasks for the dataset.
  • Shared task setup: Across all tasks, each structure contains a surface and adsorbate, with a fully periodic unit cell and at least 20 Å of vacuum along z.Initial structures are heuristically determined, and ground-truth data is computed using DFT.
  • Structure to Energy and Forces (S2EF): S2EF predicts DFT-calculated energy and per-atom forces from atomic positions, with energy referring to adsorption energy unless otherwise noted.Adsorption energy is defined from the combined surface–adsorbate system relative to the relaxed slab and relaxed gas-phase molecule.
  • Initial Structure to Relaxed Structure (IS2RS): IS2RS predicts final relaxed atomic positions from an initial structure, replacing iterative DFT force calculations that typically require hundreds of evaluations.Traditional relaxation updates atomic positions using estimated DFT forces until convergence.
  • Initial Structure to Relaxed Energy (IS2RE): IS2RE predicts the energy of the relaxed structure from its initial structure, either through iterative S2EF-based relaxation or direct energy regression.Relaxed energies are commonly used because they correlate with catalyst activity and selectivity and parameterize detailed microkinetic models.

The OC20 Dataset · 1) adsorbate selection,

OC20 provides large-scale training and evaluation data for DFT approximation and structure-relaxation tasks, spanning diverse surfaces and adsorbates. Its adsorbates are sampled from 82 renewable-energy-relevant molecules, while the dataset includes generated structures, relaxations, and additional calculations.

  • The OC20 Dataset: 640,081 relaxations are provided for training across a wide variety of surfaces and adsorbates, with intermediate structures, energies, and forces retained.The dataset supports three previously defined tasks involving DFT approximation and structure relaxation.
  • Dataset Generation: The dataset is constructed through adsorbate selection, surface selection, initial structure generation, and structure relaxation, with configuration-generation code released publicly.These four stages are described sequentially in the dataset-generation methodology.
  • Adsorbate Selection: 82 adsorbate molecules are sampled randomly, covering oxygen- or hydrogen-only, C1, C2, and nitrogen-containing chemistries selected for renewable-energy applications.Oxygen and hydrogen molecules represent water-solvated electrochemical reactions; C1 and C2 molecules support solar-fuel synthesis.
  • Surface Selection: Surface compositions are sampled with 5% unary, 65% binary, and 30% ternary probabilities, then stable bulk materials are selected from 11,451 Materials Project entries.Binary and ternary materials receive greater emphasis because they contain a wider variety of understudied materials.
  • Initial Structure Generation: Selected adsorbates are placed on selected surfaces using CatKit and ASE, with surface atoms and adsorbate binding atoms identified for initial-structure construction.Surface identification uses position, z-distance within 2 Å, under-coordination, and pymatgen Voronoi coordination environments.
  • Structure Relaxation: Relaxations use VASP until all per-atom forces are below 0.03 eV/Å, while subsurface atoms remain fixed to approximate a semi-infinite bulk condition.Calculations could run for up to 144 hours on 12 cores; intermediate structures, energies, and forces were stored.
  • MD and Rattled Calculations: Additional training data include 900 K partial molecular dynamics and rattled structures, producing approximately 950 thousand MD and 30 million rattled calculations.MD trajectories lasted 80 fs or 320 fs, while rattling used normally distributed displacements with µ = 0 and σ = 0.05.

Baseline GNN Models

The study benchmarks three state-of-the-art graph neural networks—CGCNN, SchNet, and DimeNet++—as baseline models for catalyst modeling tasks. These models operate on atomistic graphs with standardized neighborhood, atom-type, parameter-size, and energy/force-loss settings.

  • Resources: Code and pretrained baseline models implemented in PyTorch Geometric are publicly available through the Open Catalyst Project.The resources support direct use of the baseline machine-learning approaches.
  • Graph representation: GNNs represent atoms as nodes and neighboring-atom relationships as edges, iteratively updating atom embeddings through message passing.Unlike traditional descriptor-based models, the neural networks learn atomic representations during message passing.
  • Models: The benchmark evaluates Crystal Graph Convolutional Neural Network (CGCNN), SchNet, and DimeNet++ as representative recent GNN methods.The authors note that the evaluated set is not comprehensive but demonstrates what is feasible with current models.
  • Implementation: 6 Å cutoff and up to 50 nearest neighbors define graph edges, with periodic boundary conditions included in distance calculations.Atoms are additionally tagged as slab, surface, or adsorbate to let loss functions emphasize free atoms over fixed atoms.
  • Implementation: 3.6 million, 7.4 million, and 1.8 million parameters are used for CGCNN, SchNet, and DimeNet++, respectively, with roughly equivalent runtimes targeted.The corresponding hidden-channel counts are 128, 1024, and 192.
  • Training objective: The baseline loss combines computed energies and forces, while IS2RE uses only the energy term with λF = 0.The force term is defined over free atoms, whereas the full loss includes empirical energy and force weights.

Experiments

Experiments define practical DFT-grounded metrics for energy/force prediction and structure relaxation, then benchmark CGCNN, SchNet, and DimeNet++ across the OC20 tasks. DimeNet++ generally performs best, but low threshold-based accuracies show that baseline models remain far from practical usefulness.

  • S2EF metrics: S2EF evaluates energy MAE, force MAE on free atoms, and EFwT, which requires both energy and maximum force errors to meet thresholds.EFwT uses ϵ = 0.02 eV for energy and α = 0.03 eV/˚A for maximum per-atom force error.
  • S2EF results: DimeNet++ performs best across most S2EF metrics, while SchNet marginally outperforms both DimeNet++ and CGCNN on EFwT.SchNet outperforms CGCNN across all reported S2EF metrics.
  • S2EF results: Energy-only and force-only training achieve the best respective energy and force MAE, whereas balanced SchNet and DimeNet++ models perform best on EFwT.All approaches perform badly on EFwT, indicating that the baseline results remain far from practical usefulness.
  • IS2RS results: DimeNet++ outperforms SchNet on IS2RS ADwT and AFbT, but neither method produces relaxed structures with forces below practical FbT thresholds.IS2RS uses S2EF force predictions to drive L-BFGS relaxations from initial structures.
  • Relaxation results: Hybrid relaxation approaches outperform direct approaches across all metrics, but predicted energies within the tight EwT threshold range from 2% to 6%.Generalization to new catalyst compositions performs better than generalization to new adsorbates, and larger datasets could significantly improve performance.
  • Scaling and challenges: The dataset is two orders of magnitude larger than previous catalyst DFT datasets, creating scale-related training challenges for atomistic machine learning models.The largest baseline models contained approximately 10 million parameters.

DFT Relaxations

DFT calculations used VASP with periodic boundary conditions and PAW pseudopotentials. Plane-wave expansions employed 350 eV kinetic-energy cut-offs, while exchange-correlation effects used GGA with the RPBE functional.

  • Computational method: DFT calculations used VASP with periodic boundary conditions and projector augmented wave pseudopotentials.These settings define the computational framework for the relaxations.
  • Computational method: External electrons were expanded in plane waves with 350 eV kinetic-energy cut-offs.The plane-wave basis was set using a 350 eV cutoff.
  • Computational method: Exchange-correlation effects were treated with the generalized gradient approximation and the revised Perdew-Burke-Ernzerhof functional.The RPBE functional was selected because of its improved description of atomic energetics.

Adsorption Energy

The adsorption-energy section defines gas-phase reference energies for adsorbates from a linear combination of molecular references and specifies a force-based exclusion criterion for most tasks.

  • Adsorption Energy: Gas-phase reference energies, Egas, were computed for each adsorbate as linear combinations of N2, H2O, CO, and H2.The resulting atomic energies are given in Table 5.
  • Adsorption Energy: Systems with fmax > 0.05 eV/˚A were excluded from all tasks except S2EF.The criterion applied to systems that converged and completed successfully.
  • Adsorption Energy: Table 5 reports per-atom energies for individual adsorbate atoms used to calculate gas-phase reference energies for adsorbate molecules.

Computational Workflow

Figure 9 illustrates the workflow used to sample from the dataset and perform calculations.

  • Figure 9 illustrates the computational workflow.
  • The workflow includes sampling from the dataset.
  • The workflow is used to perform calculations.

Graph Construction

The method represents periodically repeated atomic structures as radius graphs that capture local 3D environments. Directed, bidirectional edges and periodic multi-edges encode neighboring atoms across repeated cells and their distinct relative distances.

  • Graph Construction: Atoms are represented as graph nodes, with edges connecting pairs within a specified cutoff distance.The construction uses a radius graph over atoms in the periodically repeated 3D unit cell.
  • Graph Construction: Each neighboring pair receives directed edges in both directions, making the graph bidirectional.An edge is drawn from atom j to atom i and vice versa whenever the cutoff condition is met.
  • Graph Construction: Periodic images can create multiple directed edges between the same atoms, with each edge representing a different repeated cell and relative distance.These edges carry distinct edge features and preserve periodic boundary conditions in the atom-centric representation.
  • Graph Construction: The directed multigraph precisely captures the local 3D structure surrounding each atom while accounting for periodic boundary conditions.The atom-centric view considers each atom as the center node individually.

Graph Pairwise Similarity

The section defines mean pairwise similarity as a graph-based measure of dataset diversity and describes its estimation across sampled systems. It also documents baseline-model adaptations and explains why OC20 IS2RE errors exceed prior studies while remaining consistent on a literature dataset.

  • Graph Pairwise Similarity: Mean pairwise similarity measures dataset graph diversity as the mean of the upper-triangle similarity-matrix elements, excluding the diagonal.The similarity matrix uses graphs and the molecular kernel from GraphDot.
  • Graph Pairwise Similarity: 1 to 0: mean pairwise similarity ranges from identical graphs to maximally dissimilar graphs and is comparable across datasets only with consistent graph and kernel parameters.For Figure 6, 1000 systems were randomly sampled from a 10,000-system subsample per dataset, with six repetitions.
  • Baseline Implementations: Periodic boundary conditions were added to the PyTorch Geometric implementations of SchNet and DimeNet++.The baseline models were implemented using PyTorch Geometric.
  • Baseline Implementations: A Gaussian basis function improved CGCNN edge features, and position gradients enabled force predictions beyond its original energy-only implementation.These changes were made to adapt CGCNN to the dataset and tasks.
  • IS2RE Results: 0.57 eV MAE: baseline models trained only on the OC20 CO subset substantially exceeded the approximately 0.190 eV MAE reported for a CGCNN-based literature-dataset model.The OC20 IS2RE task is more challenging because its dataset is larger, more diverse, sparser, and more uniformly sampled, and predicts final energy from the initial structure.

Adsorbates Included

The dataset includes four monatomic species and common intermediates relevant to renewable-energy challenges, but it is not a comprehensive catalog of adsorbates. Most adsorbates were initialized as mono-dentate, while larger molecules known to bind bi-dentately were initialized in those configurations.

  • Adsorbates Included: Four monatomic species and common intermediates for renewable energy challenges define the included adsorbate list.The full list is provided in Table 10.
  • Adsorbates Included: The adsorbate set is not comprehensive because organic molecules have combinatorially many possible structures and larger molecules have even more configurations.Larger molecules such as C3 are also relevant but have an even larger number of possible configurations.
  • Adsorbates Included: Most adsorbates were mono-dentate, whereas larger molecules known to bind bi-dentately were initialized in bi-dentate configurations.The atoms considered for each adsorption location are indicated by an asterisk.

Train/Test/Validation Splits · Tight Binding Baseline

The dataset uses adsorbate-based validation and test subsplits, while a tight-binding baseline shows promising speed but important periodic-boundary limitations. Additional off-equilibrium structures from rattling and molecular dynamics improve data diversity and sample efficiency for force prediction.

  • Train/Test/Validation Splits: Validation subsplits reserve seven specified adsorbates, including *CH, *CHO, *COH, and *NH2.Asterisks denote binding atoms.
  • Train/Test/Validation Splits: Test subsplits reserve seven different adsorbates, including *CH2*CH2, *CO, *COHCH2, and *ONNO2.The reserved test adsorbates are distinct from those assigned to validation.
  • Tight Binding Baseline: 100 random validation systems were evaluated with xTB using GFN0 parameters through the ASE interface to assess energies, forces, and relaxed structures.The baseline targets catalysis calculations at lower computational cost than DFT.
  • Tight Binding Baseline: xTB calculations were fast, but the code’s non-periodic design made incorporation of periodic boundary conditions an ongoing challenge.The reported limitation affects the use of tight binding for periodic catalytic systems.
  • Additional Data: Rattled & Molecular Dynamics: Off-equilibrium structures were added through structural perturbations and molecular dynamics to diversify the dataset.These approaches generate additional structures beyond the relaxation data.
  • Additional Data: Rattled & Molecular Dynamics: Approximately 30 million rattled single-point calculations yielded 17M filtered S2EF training data points.Rattled structures were sampled from relaxation pathways, randomly displaced with µ = 0 and σ = 0.05, and evaluated with DFT.
  • Additional Data: Rattled & Molecular Dynamics: DimeNet++ outperforms SchNet on force-prediction metrics, with lower Force MAE, higher Force cosine, and higher AFbT.The comparison is reported for S2EF and IS2RS results.
  • Additional Data: Rattled & Molecular Dynamics: Both models trained on S2EF-20M + MD with 58M samples outperform corresponding S2EF-All models with 134M samples on AFbT.The result indicates that molecular-dynamics data provides a complementary learning signal and improves sample efficiency.

Results on Validation splits · Changelog

Validation results are reported for the S2EF, IS2RS, and IS2RE tasks in Tables 13, 14, and 12, while subsequent changelog entries document model replacements, metric additions, dataset corrections, and metric re-evaluations.

  • Results on Validation splits: Full validation-split results are provided for S2EF, IS2RS, and IS2RE in Tables 13, 14, and 12, respectively.The tables cover energy-and-force prediction, relaxed-structure prediction, and relaxed-state energy prediction.
  • Results on Validation splits: Table 12 evaluates IS2RE using energy MAE and Energies within a Threshold (EwT) for models trained on the All training dataset.The table reports relaxed-state energy prediction from the initial structure.
  • Results on Validation splits: Table 13 evaluates S2EF using energy MAE, force MAE, force cosine, and Energies and Forces within Threshold (EFwT) for S2EF-All models.The models are trained on the entire training dataset.
  • Results on Validation splits: Table 14 evaluates IS2RS with Average Distance within Threshold (ADwT), using S2EF models trained on the All training dataset and the initial structure as an IS baseline.All values are percentages, higher is better; FbT and AFbT are computed only for test splits.
  • Changelog: DimeNet90 results were replaced with DimeNet++, which is described as more memory-efficient and slightly better-performing.The changelog identifies this as a document update.
  • Changelog: Force cosine similarity was added as an S2EF metric because it correlates better with downstream IS2RS AFbT.Rattled/MD data experiments were also added.
  • Changelog: 81 systems were removed after convergence issues were discovered, representing approximately 0.00675% of the data; models were not retrained because the affected amount was negligible.Some systems were later removed or modified for convergence and trajectory-stitching issues, again without retraining S2EF models.
  • Changelog: An IS2RE via relaxation bug was resolved, updated metrics outperformed direct-based approaches, and validation/test removals affected approximately 0.167% of IS2RE and 0.043% of S2EF.S2EF, IS2RE, and IS2RS metrics were re-evaluated with the updated splits; IS2RE models were retrained and IS2RS ADwT metrics re-evaluated.
Loading 2010.09990v5…