~/aixsci
200 records · all checked

materials-chemistry/ai produced the result/The Journal of Physical Chemistry C 2023 · v2

Neural network potentials trained by repeatedly checking where models disagree on hydrogen sticking to copper

Researchers built fast neural-network models of the forces between hydrogen molecules and copper surfaces, growing the training data by letting committees of models flag the configurations they disagreed about and labelling those with quantum calculations.

1. Build initial reference dataset2. Train MLIP committees and optimise hyperparameters3. Run reactive scattering molecular dynamics4. Flag high-uncertainty structures during dynamics5. Cluster and sparsify selected structures6. Label new structures with DFT and extend dataset7. Compare simulated observables with reference data

spectrum · one line per step, placed by what the step does · bright lines used AI

Machine Learning Interatomic Potentials for Reactive Hydrogen Dynamics at Metal Surfaces Based on Iterative Refinement of Reaction Probabilities
The Journal of Physical Chemistry C, 2023

doi:10.1021/acs.jpcc.3c06648 · record aix-00176 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Simulation surrogate, Experimental design
Model family
Graph neural network, Clustering
Checked by
Benchmark
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Schematic showing DFT data training initial models, running gas-surface scattering dynamics to select new data points, and iterating until observables converge.
Schematic of the iterative adaptive sampling procedure refining machine learning potentials via DFT-trained scattering simulations.Figure 2 from Stark et al., The Journal of Physical Chemistry C 2023 · source · CC BY · resized

When a hydrogen molecule strikes a metal surface, it may bounce off or split apart and stick. Chemists want to predict how often each happens, because that step underlies how metals take up hydrogen and how catalysts work. Predicting it means following the molecule's motion step by step, and at every step you need the forces acting on each atom. Those forces come from density functional theory, a quantum method for working out the energy of a set of atoms. It is accurate but slow, and a single reaction probability needs enormous numbers of simulated collisions. The cost is what blocks the calculation, not the theory.

The researchers set out to replace the quantum step with a trained model that returns the same energies and forces far more cheaply, for hydrogen on four differently cut copper surfaces, and to see how two model designs compared.

Where AI came in

Two neural networks, SchNet and PaiNN, were trained from scratch to predict energies and forces for copper slabs with and without hydrogen. They acted as the force engine inside the collision simulations, standing in for the quantum calculation at every step; the reported sticking probabilities, each drawn from 10,000 simulated collisions, come from running them.

The models also chose their own training data. Three copies of a network, trained on different splits of the data, were run together and their predictions compared every fifth step of a simulation. Where they disagreed beyond a set margin, that configuration was collected. Those configurations were grouped by similarity using k-means clustering, screened by hand, then labelled with quantum calculations and added to the training set. Four rounds of this grew the set from 2530 to 4230 data points.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Schematic showing DFT data training initial models, running gas-surface scattering dynamics to select new data points, and iterating until observables converge.
Schematic of the iterative adaptive sampling procedure refining machine learning potentials via DFT-trained scattering simulations.Figure 2 from Stark et al., The Journal of Physical Chemistry C 2023 · source · CC BY · resized

The study builds machine-learning interatomic potentials for molecular hydrogen scattering on four copper facets — Cu(111), Cu(100), Cu(110) and Cu(211) — using an adaptive sampling loop in which committees of three models flag configurations they disagree about during scattering dynamics, those configurations are clustered and labelled with SRP48 density functional theory, and the models are retrained. Four iterations grew the training set from 2530 to 4230 data points, and sticking probabilities were obtained from 10,000 trajectories per model, facet, rovibrational state and collision energy. The invariant SchNet model required four iterations before its ground-state sticking probabilities approached published quasi-classical reference data, and its potential energy surface cuts remained non-smooth; the equivariant PaiNN model, trained on the same datasets, matched the reference probabilities closely from the initial dataset and gave smooth energy landscapes. For the best final models the paper reports PaiNN energy errors three times lower in MAE and more than four times lower in RMSE, and force MAE and RMSE more than five times lower, than SchNet.

How AI was used

Two message-passing neural network interatomic potentials, SchNet (SchNetPack v1.0.0) and PaiNN (SchNetPack dev branch), were trained from scratch on SRP48 DFT total energies and forces for 6-layer copper slabs with and without hydrogen, with a combined energy–force loss (final weights 0.05 and 0.95) and hyperparameters chosen by five-fold cross-validation, settling on 7 interaction blocks, 512 features and a 4 Å cutoff. The initial dataset was assembled from ab initio molecular dynamics in FHI-aims plus H2 scattering trajectories run with a SchNet potential trained on an externally supplied dataset. Training data were then grown adaptively: reactive scattering molecular dynamics in NQCDynamics.jl was run with a committee of three models differing in train/test split, the standard deviation between their energy predictions was evaluated every fifth step, and structures exceeding a manual threshold (0.03 eV, then 0.025 eV) were collected, described by inverse H–H and H–surface distances, reduced by principal component analysis, grouped with k-means clustering in scikit-learn, manually screened, labelled with DFT and added to the dataset before retraining. The same model committees were run as the force engine for the scattering trajectories from which sticking probabilities and their epistemic uncertainties were computed, and for potential energy surface cuts and minimum energy path calculations.

The shape of the work

Structural · the record, drawn

ACQUISITIONTRAININGSIMULATIONSCREENINGPREPARATIONACQUISITIONVALIDATION1234567AIAIAIAIAIBuild initialreference datasetTrain MLIPcommittees andoptimise hyperpa…Run reactivescatteringmolecular dynami…Flaghigh-uncertaintystructures durin…Cluster andsparsify selectedstructuresLabel newstructures withDFT and extend d…Compare simulatedobservables withreference data↤ simulation↤ simulation↤ simulation↤ expert judgement↤ manual curationloops back · 4 rounds
AI stepNo AI↤ what the AI stood in for
1Acquisition
AI

Build initial reference dataset

Obtaining raw data, whether by measurement, download or retrieval. The AI stood in for simulation.

The initial MD simulations were performed with a SchNet MLIP with standard settingswhere the paper describes this · verbatim
in the paper
2Training
AI

Train MLIP committees and optimise hyperparameters

Fitting model parameters, including fine-tuning an existing model. The AI stood in for simulation.

After the fourth adaptive sampling iteration, we performed a detailed parameter optimization for both SchNet and PaiNN modelswhere the paper describes this · verbatim
in the paper
3Simulation
AI

Run reactive scattering molecular dynamics

Numerical or physics simulation, including where a learned surrogate replaces it. The AI stood in for simulation.

Sticking probabilities were calculated using data extracted from 10,000 trajectories for every model, surface facet, rovibrational state, and initial collision energy reported.where the paper describes this · verbatim
in the paper
4Screening
AI

Flag high-uncertainty structures during dynamics

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for expert judgement.

at every fifth step of the MD simulation, we evaluate the potential using 3 models trained with different train/test splitswhere the paper describes this · verbatim
in the paper
5Preparation
AI

Cluster and sparsify selected structures

Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for manual curation.

the dataset is analyzed using k-means clusteringwhere the paper describes this · verbatim
in the paper
6Acquisition
no AI

Label new structures with DFT and extend dataset

Obtaining raw data, whether by measurement, download or retrieval. Its result feeds back into an earlier step.

The adaptive sampling loop was iterated 4 times until we obtained a satisfactory level of accuracywhere the paper describes this · verbatim
in the paper
7Validation
no AI

Compare simulated observables with reference data

Testing outputs against ground truth.

We compare our simulation results to the literature data based on quasi-classical dynamics (QCD) simulations with a CRP potentialwhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The reported dynamic observables (sticking probabilities from 10,000-trajectory ensembles) are produced by running the trained SchNet and PaiNN interatomic potentials; the study's conclusions about the two architectures depend entirely on them.

+What the AI was for
we employ ensemble learning to adaptively generate training data while assessing model performance with full uncertainty quantificationwhere the paper describes this · verbatim
+Model families
+How it was taught
SupervisedActive learningUnsupervisedin the paper
+Models named
SchNet SchNetPack v1.0.0 master branch · Trained from scratchPaiNN SchNetPack developmental ("dev") branch, base for v2.0 · Trained from scratchk-means clustering (scikit-learn) · Trained from scratchin the paper
+How results were checked
Benchmarkin the paper
we finally investigated how general and robust the models are by comparing them against experimental sticking probabilitieswhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The dynamics for high-error structure search and the clustering scripts are available in the GitHub repositorywhere the paper describes this · verbatim
+Compute
High-performance computing via the University of Warwick Scientific Computing Research Technology Platform, ARCHER2 and HPC Midlands+; the models are reported as roughly 10 times faster than DFT. No accelerator hours or wall-clock training times given.in the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 3 items
  • Trained model weightsWhether the trained model is available is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of k-means clustering (scikit-learn)Which version of the model was used is not stated.

About this article

Record aix-00176, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error